Skip to main content
U.S. flag

An official website of the United States government

Official websites use .gov
A .gov website belongs to an official government organization in the United States.

Secure .gov websites use HTTPS
A lock ( ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.

The TEVV-Athlon Framework for Evaluating AI Systems

Input Sought on Initial Public Draft of NIST AI 200-2 through October 6, 2026

Date of announcement: August 7, 2026

Artificial intelligence (AI) applications span a wide range of disciplines and use cases. Within these fields, requirements and evaluation methods can vary significantly, with many evaluation components tailored to specific applications. To build trust and encourage adoption, stakeholders can conduct a thorough assessment of AI systems. The NIST AI Risk Management Framework specifically calls for a Test, Evaluation, Verification, and Validation (TEVV) methodology. This report introduces the TEVV-Athlon framework, a structured approach for assessing the real-world impact and outcomes of AI systems. The aim of the TEVV-Athlon framework is to be extensible, adaptable, and customizable to accommodate the wide variety of AI applications. This includes statistical machine learning models, large language models, multi-modal models, agentic systems, and many other types of AI technologies.

Abstract

Test, evaluation, verification, and validation (TEVV) of artificial intelligence (AI) systems is used to provide evidence that systems can effectively meet individual or organizational goals while minimizing negative impacts. Due to the extensive variety of AI use, a flexible approach is needed for organizations to develop and conduct measurement approaches that are customized to their needs. This paper introduces the TEVV-Athlon Framework, a four-stage method for developing customized assessments of AI systems based on organizational TEVV objectives. The framework produces a TEVV-Athlon, an assessment where AI systems are tested via a set of Events and Tools which produce data on Blocks related to measurement concepts of interest. To illustrate the approach, an example TEVV-Athlon is constructed using the framework. Practical guidance for conducting a TEVV-Athlon is also provided. The TEVV-Athlon Framework can be applied as needed to produce meaningful information about AI system performance and help organizations measure the impact of their AI systems.

Request for Input

NIST invites input on any aspect of this draft document, particularly:

  1. Definitions and uses of the terms AI testing, evaluation, validation, and verification (TEVV), as well as other AI measurement terms or sources that inform these definitions.
  2. The flexibility, scope, and applicability of the TEVV-Athlon framework across different AI TEVV processes, activities, systems, and contexts.
  3. Types of AI TEVV processes or activities that may not be adequately addressed by the current framework.
  4. The usefulness of the TEVV-Athlon framework for developing new AI TEVV processes or activities and TEVV for novel or emerging AI systems.
  5. Aspects of the TEVV-Athlon framework that may warrant clarification, revision, removal, or expansion.
  6. Additional concepts, provisions, examples, or supporting materials that could improve the clarity, completeness, or practical utility of the TEVV-Athlon framework.
     

A 60-day comment period opened August 7, 2026, and closes on October 6, 2026. Feedback can be emailed to TEVV-Athlon [at] nist.gov (TEVV-Athlon[at]nist[dot]gov) with “NIST AI 200-2” in the subject line. Comments may be sent via email, or email attachment in any of the following unlocked formats: HTML; ASCII; Word; RTF; Excel; or PDF.

NIST may release comments. All comments are subject to release under the Freedom of Information Act (FOIA). NIST requests that comments do NOT contain proprietary information.

NIST encourages all stakeholders to provide input, including organizations with experience conducting AI evaluations as well as users of AI evaluation reports – for instance, business decision-makers, procurement specialists, researchers, and technical staff.

NIST researchers and staff may use a variety of software tools to help summarize or analyze your comments, including AI. If AI is used, your data will not be used to train AI models.

The document is available here: https://doi.org/10.6028/NIST.AI.200-2.ipd

Contacts

Created August 4, 2026, Updated August 5, 2026
Was this page helpful?