Skip to main content
U.S. flag

An official website of the United States government

Official websites use .gov
A .gov website belongs to an official government organization in the United States.

Secure .gov websites use HTTPS
A lock ( ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.

Guardians of Forensic Evidence

Summary

Researchers from NIST are developing the Forensics Deepfake Evaluation program. America's AI Action Plan specifically cites this project, recommending that NIST consider developing it into a formal guideline and voluntary forensic benchmark. This program is designed to advance the development of forensic technologies for automatically detecting deepfakes and AI-generated media. Its goals include fostering research, establishing a reference baseline detection system, supporting the transition from lab prototypes to real-world products, and enhancing detection tool generalization. The evaluation follows a seven-step process, starting with engaging stakeholders and defining the program, followed by developing a framework, collaborating with experts, running evaluations with participants, reporting findings, and ultimately upgrading and iterating the program. 

Description

The program addresses critical challenges including a gap between high research accuracy and a lack of ease-of-use in real-world applications, as well as the need for improved generalization capability and robustness against post-processing and anti-forensics filters. The specific image deepfake detection task involves studying these capabilities by using data from sources like StyleGANs and Stable Diffusion. Detectors are trained on older deepfake generation methods and tested against both older and newer techniques, such as those that have undergone Gaussian blur or video compression, to evaluate their resilience. To accomplish the program's goals, NIST is pursuing the following efforts:

Request For Information

NIST is issuing a Request for Information (RFI) on the Federal Register. The goal of the RFI is to gather data on how forensic examiners perform their duties: current tools, evidence types, and processes in use, as well as the availability and application of AI analysis tools within forensic shops. A link will be available here upon publication of the RFI.

Best Practices & Guidelines

NIST will develop a framework to emphasize training forensic examiners to function as active software validators capable of independently executing specialized, scenario-specific evaluations to address operational casework complexities. It will provide suggested operational protocols and methods to ensure evaluation integrity, including guidelines related to data processing, training and testing. The core methodology utilizes uniform validation protocols to stress-test analytic systems against real-world deepfakes generated by emergent architectures. Key components of these guidelines focus on:

  • Defining Evaluation Tasks: Establishing targeted forensic questions for validation, such as image authenticity detection, person identity verification (checking face-swaps), image manipulation localization, source verification, or provenance reconstruction.
  • Curation of Forensically Relevant Data: Collecting independent reference sets that mirror real-world forensic casework. Data should stress-test generalizability by training detectors on older fakes and testing on both older and newer generation techniques. The guidelines require data representative of use conditions, content, and generators , and the use of "dirty," "post-processed" evidence, such as low-bitrate surveillance footage and media subjected to compression artifacts typical of social media redistribution.
  • Measurement and Continuous Assessment: Performance analysis relies on standard statistical methods, specifically Receiver Operating Characteristic (ROC) curves and Area Under the ROC Curve (AUC) metrics, to provide threshold-independent summaries of classification capability. Because generative threat landscapes are dynamic, guidelines mandate continuous validation lifecycles with periodic reassessment of tools against novel threats and regression testing following any software updates.

Challenge Kit

To bridge the gap between research-grade performance and operational forensic use, this initiative introduces the Deepfake Challenge Kit. Rather than a rigid, final test, this kit serves as a hands-on educational sample package designed to teach the forensic community how to perform systematic tool validation and evaluation. The Challenge Kit acts as the practical training vehicle for a completely new approach in forensic testing. By working through this sample package, examiners learn how to assess a tool's performance under realistic forensic conditions using three integrated components: 

  • Preliminary Study Dataset:  A curated sample collection of authentic and manipulated media that simulates realistic forensic evidence conditions, serving as a training ground for examiners to build their own agency-specific test sets.
  • Scoring Package: A standardized analytical toolset used to calculate performance metrics, teaching examiners how to run and interpret Receiver Operating Characteristic (ROC) analysis and Area Under the Curve (AUC).
  • Technical Framework Document: This comprehensive technical framework document guides the entire validation and evaluation lifecycle, including task definition, forensic dataset curation, tool input/output standardization, and the generation of scientifically rigorous performance reports suitable for judicial scrutiny.
     
Created May 21, 2026, Updated August 17, 2026
Was this page helpful?