Researchers from NIST are developing the Forensics Deepfake Evaluation program. America's AI Action Plan specifically cites this project, recommending that NIST consider developing it into a formal guideline and voluntary forensic benchmark. This program is designed to advance the development of forensic technologies for automatically detecting deepfakes and AI-generated media. Its goals include fostering research, establishing a reference baseline detection system, supporting the transition from lab prototypes to real-world products, and enhancing detection tool generalization. The evaluation follows a seven-step process, starting with engaging stakeholders and defining the program, followed by developing a framework, collaborating with experts, running evaluations with participants, reporting findings, and ultimately upgrading and iterating the program.
The program addresses critical challenges including a gap between high research accuracy and a lack of ease-of-use in real-world applications, as well as the need for improved generalization capability and robustness against post-processing and anti-forensics filters. The specific image deepfake detection task involves studying these capabilities by using data from sources like StyleGANs and Stable Diffusion. Detectors are trained on older deepfake generation methods and tested against both older and newer techniques, such as those that have undergone Gaussian blur or video compression, to evaluate their resilience. To accomplish the program's goals, NIST is pursuing the following efforts:
NIST is issuing a Request for Information (RFI) on the Federal Register. The goal of the RFI is to gather data on how forensic examiners perform their duties: current tools, evidence types, and processes in use, as well as the availability and application of AI analysis tools within forensic shops. A link will be available here upon publication of the RFI.
NIST will develop a framework to emphasize training forensic examiners to function as active software validators capable of independently executing specialized, scenario-specific evaluations to address operational casework complexities. It will provide suggested operational protocols and methods to ensure evaluation integrity, including guidelines related to data processing, training and testing. The core methodology utilizes uniform validation protocols to stress-test analytic systems against real-world deepfakes generated by emergent architectures. Key components of these guidelines focus on:
To bridge the gap between research-grade performance and operational forensic use, this initiative introduces the Deepfake Challenge Kit. Rather than a rigid, final test, this kit serves as a hands-on educational sample package designed to teach the forensic community how to perform systematic tool validation and evaluation. The Challenge Kit acts as the practical training vehicle for a completely new approach in forensic testing. By working through this sample package, examiners learn how to assess a tool's performance under realistic forensic conditions using three integrated components: