Skip to main content
U.S. flag

An official website of the United States government

Official websites use .gov
A .gov website belongs to an official government organization in the United States.

Secure .gov websites use HTTPS
A lock ( ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.

Biosecurity for Synthetic Nucleic Acid Sequences

Summary

Balancing the growth of the bioeconomy with the inherent risks associated with the potential misuse of artificial intelligence (AI) related to nucleic acid synthesis will require comprehensive and sustainable nucleic acid synthesis screening practices and risk mitigation strategies including standards, databases, tools, and capacities to identify, track, and defend against emerging sequences of concern (SOCs).

NIST is partnering with key stakeholders to 1) improve current screening standards and practices, and 2) mitigate emerging risks particularly through the use of AI biodesign tools. This effort is supported by NIST biometrology, engineering biology, and cybersecurity capabilities.

Description

Genetic codes
Credit: Sergei Drozd/Shutterstock

Synthetic nucleic acid technologies are fundamental to U.S. biotechnology and biomanufacturing innovation. However, like other transformative technologies, they carry dual-use potential: while they can drive significant progress, they pose risks of unintentional or deliberate misuse to engineer harmful biological systems. With the increased convergence of biotechnology and AI, the possibility also exists that AI could be used to design entirely novel DNA sequences, undetectable by current sequence screening tools, that may increase harm.

NIST has been engaging with industry and relevant stakeholders to develop comprehensive, scalable, and verifiable synthetic nucleic acid procurement screening mechanisms for commercial synthetic DNA manufacturers. NIST is also leveraging its biotechnology program to support emerging measurement needs and challenges associated with safety and security in the synthesis of nucleic acids.

Selected Programs and Accomplishments

Improving Current Screening Practices

NIST has developed a benchmark dataset consisting of 200 bp sequences with known performance metrics for testing of providers' baseline sequence screening capabilities.  The dataset was tested by six screening tool developers and was deemed to be fit-for-purpose as described in this manuscript.  NIST is currently developing a revised dataset to address advancing nucleic acid screening guidelines, including a reduction in sequence length to 50 bp.

Left, “attestation dataset”, a series of overlapping circles similar to a venn diagram with different colors. There is a central line dissecting the horizontal middle of the outermost circle the top corresponding to ‘threat’ and the bottom corresponding to ‘safe’. The outermost circle contains a label ‘200 bp fragments’ and within that has ’pathogens and ‘non-pathogens’ and within pathogens, “TP”. In the ‘safe’ zone below the line there are “BSATs” and “TN”. To the right of this image is a circular diagram
Iterative process for generating benchmark datasets with known performance metrics.

Working with broader stakeholders, NIST completed a Draft Standard Guide for Nucleic Acid Providers that harmonizes nucleic acid sequence approaches and standardizes data to enable interoperability and integration.  (See Annex III in this EBRC Report.)

Through a grant to the Engineering Biology Research Consortium (EBRC), NIST and EBRC held six virtual and one two-day, in-person workshops aimed at developing robust tools, capabilities, and standards. 
Workshop report can be found here

NIST also contributed to recent ISO standards and ensured alignment of requirements for biosafety and biosecurity with current best practices.
ISO 20688-1:2020 focuses on synthesized oligonucleotides
ISO 20688-2:2024 focuses on synthesized gene fragments, genes, and genomes
• ISO 20688-3 (tentative) focused on providing standardized sequence screening is under consideration.

Monthly Baseline Screening

In August 2025, NIST implemented monthly testing to support comprehensive, scalable, and verifiable synthetic nucleic acid procurement screening mechanisms.  Through CRADAs with IBBIS, MITRE, and SecureDNA, NIST provides a dataset to each CRADA partner that consists of 1000 sequences:  200 true positives, 200 true negatives, and 600 ungraded sequences.  Nucleic acid providers can obtain the test set, label sequences as "flag" or "no flag," and return their results for scoring.  Anonymized results are then sent to NIST for ongoing analysis.

As of July 2026, provider results showed a median sensitivity of 0.9675 and median accuracy of 0.9788.  Both metrics achieve the current thresholds for scoring pass (sensitivity >0.95 and accuracy >0.75).  In addition, provider and non-governmental agency (NGO) feedback helps shape future months' datasets, and NIST continuously removes sequences that are not functionally concerning (i.e., false positives).

Stress Testing

NIST regularly orders DNA-based materials composed of viral genomic fragments as a part of its ongoing efforts to provide reference materials and standards that support validation and development of assays for detecting emerging and re-emerging pathogens.  In June 2025, NIST placed an order for three plasmids, each designed to include PCR target regions for detecting and discriminating Mpox (clades I and II), variola virus, and all Orthopoxviruses more broadly.  When NIST Placed this order with a single provider, the provider flagged the order for the presence of variola virus sequence and required NIST to complete follow-up customer screening.  NIST subsequently submitted the same order to additional providers, either as orders for plasmids, dsDNA fragments, or both.

Sankey diagram showing intermediate screening and final results for twelve synthetic nucleic acid sequence orders containing viral sequences
Synthetic NA ordering process differs among providers.  Colors indicate initial order processes (blue), technical issues (yellow), flagging / biosecurity issues (orange), fulfilled orders (green), partial fulfillment (olive), cancelled orders (grey).

Of the twelve orders, three were processed by the provider without follow-up because either the provider did not screen for SOCs; the provider conducted sequence screening that identified the sequence(s) as safe; or the provider's screening identified a SOC, but the provider recognized NIST as a legitimate customer.  For the remaining nine orders, the provider conducted some degree of follow-up with NIST before providing the sequences.

This exercise afforded useful insights into DNA providers' technical capabilities, sequence and customer screening, and the overall process for interpreting and implementing biosecurity requirements in the current, fragmented, global regulatory landscape.

Mitigating Emerging Risks
Original proteins have 100% of native activity, but it is unknown what activity will be exhibited by AI-generated protein sequences with varying similarity constraints to the original proteins.  In our work, we have tested the activity of synthetic homologs generated, shown in gray, against what we consider basic, moderate, and advanced AI-design challenges starting with known original protein sequences and structures, in yellow.
Generation of synthetic homologs of wildtype protein targets to assess synthetic protein activities.

NIST has leveraged its predictive Engineering Biology program to conduct one of the first large-scale experimental validations of AI-generated protein sequences using safe proteins as SOC proxies, as described in this preprint manuscript.  Working with Microsoft and Twist Biosciences, we designed, basic, moderate, and advanced "difficulty classes" that considered protein expression, folding, and interaction.  Results showed that AI biodesign tools generated synthetic homologs with predicted structure similar to native templates without necessarily retaining function.  This study provided insights into the real-world predictive performance of AI biodesign tools.

More recently, in a collaboration with Harvard Medical School and the Align Foundation, NIST conducted a larger-scale evaluation of generative AI capabilities by measuring the activity of over 30,000 enzyme variants that were generated using seven different generative models spanning multiple data modalities and model classes, as reported here.  The results show that different generative AI approaches give different levels of performance, with structure-based models performing the best overall.

NIST, Profluent Bio, and the Align Foundation have demonstrated that these very-large-scale evaluations can be routinely performed in a timely manner that is relevant to stakeholders in the private sector and biosecurity communities.  Together, we produced a dataset that included activity measurements for approximately 40,000 enzyme variants generated using Profluent's AI-driven protein design capabilities.  The total turn-around time was just over two months, from receipt of the AI-generated sequences from Profluent to the delivery of the fully processed dataset.

Leveraging the NIST Cybersecurity and Privacy of Genomic Data effort and the testbeds under development, NIST will test and validate secure data transfer of sequence screening.  The transmission process will be documented and areas of cybersecurity requirements against the draft Genomic CSF will be highlighted.  NIST is also exploring opportunities to apply Cybersecurity Supply Chain Risk Management and SP 800-63 principles to support due diligence and customer verification.

Any mention of commercial products within NIST web pages is for information only;
it does not imply recommendation or endorsement by NIST.

Created March 25, 2024, Updated August 3, 2026
Was this page helpful?