Skip to main content
U.S. flag

An official website of the United States government

Official websites use .gov
A .gov website belongs to an official government organization in the United States.

Secure .gov websites use HTTPS
A lock ( ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.

Search Publications

Search Title, Abstract, Conference, Citation, Keyword or Author
Published Date
Displaying 1 - 25 of 36

Photopolymer Additive Manufacturing 2025 Workshop Report: Building a Unified Vision from Research to Regulation

June 10, 2026
Author(s)
Callie Higgins, Jason Killgore, Mike Idacavage, Vince Anewenter, Mickey Fortune, Gary Cohen, Perri Katzman, Jessica Hemond, Spencer Loveless, Michael Gould
The third biannual Photopolymer Additive Manufacturing Alliance Workshop was held on September 15-16, 2025, at the University of Colorado Boulder to continue its mission of advancing photopolymer additive manufacturing (PAM). Building on the 2023 PAMA

The 34th Text REtrieval Conference (TREC 2025)

March 24, 2026
Author(s)
Ian Soboroff, George Awad
TREC 2025 is the thirty-fourth edition of the Text REtrieval Conference (TREC). The main goal of TREC is to create the evaluation infrastructure required for large-scale testing of information retrieval (IR) technology. This includes research on best

Expanding the AI Evaluation Toolbox with Statistical Models

February 17, 2026
Author(s)
Andrew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita Rao, Julia Sharp, Amanda Bergman
Benchmarks are widely used to evaluate and compare the performance of artificial intelligence systems. However, some approaches to computing benchmark metrics produce invalid uncertainty estimates or make unrecognized assumptions about the evaluation

Assessing Risks and Impacts of AI (ARIA): Pilot Evaluation Report

November 13, 2025
Author(s)
Razvan Amironesei, Afzal Godil, Craig Greenberg, Kristen Greene, Johnston Patrick Hall, Theodore Jensen, Jonathan Fiscus, Noah Schulman
This document describes the procedure used for a pilot of NIST's Assessing Risks and Impacts of AI (ARIA) evaluation: ARIA 0.1. Five organizations participated, submitting a total of 7 AI applications to be evaluated. In this document, we first describe

From Traditional Topic Models to LLM Topic Models: Can Large Language Models Replace Traditional Topic Models?

August 1, 2025
Author(s)
Zongxia Li, Lorena Calvo Bartolome, Alexander Hoyle, Daniel Stephens, Paiheng Xu, Alden Dima, Jordan Boyd-Graber, Juan Fung
A common use of NLP is to facilitate the understanding of large document collections, with models based on Large Language Models (LLMs) replacing probabilistic topic models. Yet the effectiveness of LLM-based approaches in real-world applications remains

2025 NIST GenAI (Pilot): Code Challenge Evaluation Plan

July 16, 2025
Author(s)
Peter Fontana, Yooyoung Lee, Hariharan Iyer, Sonika Sharma
We are launching a pilot for measuring and evaluating unit tests generated by Artificial Intelligence (AI) for testing elementary python code. This pilot will provide an environment that will facilitate the development and improvement of the abilities of

Quasi-Deterministic Channel Propagation Model for Human Sensing: Gesture Recognition Use Case

July 9, 2025
Author(s)
Jack Chuang, Raied Caromi, Jelena Senic, Samuel Berweger, Neeraj Varshney, Jian Wang, Anuraag Bodi, Camillo Gentile, Nada Golmie
We describe a quasi-determinstic channel propagation model for human gesture recognition reduced from real-time measurements with our context aware channel sounder, considering four human subjects and 20 distinct body motions, for a total of 120,000

2024 NIST GenAI (Pilot Study): Text-to-Text Evaluation Overview and Results

June 25, 2025
Author(s)
Hariharan Iyer, Seungmin Seo, Lukas Diduch, Kay Peterson, George Awad, Yooyoung Lee
The 2024 NIST Generative AI (GenAI) Pilot Study focuses on evaluating text-to-text (T2T) generation and discrimination tasks to assess the capabilities and limitations of generative AI models and AI detectors. The study aims to measure the effectiveness of

Experimental Evaluation of AI-Driven Protein Design Risks Using Safe Biological Proxies

June 20, 2025
Author(s)
Svetlana Ikonomova, Bruce Wittmann, Fernanda Piorino Macruz de Oliveira, David Ross, Samuel Schaffter, Olga Vasilyeva, Elizabeth Strychalski, Eric Horvitz, James Diggans, Sheng Lin-Gibson, Geoffrey Taghon
Advances in machine learning are providing new abilities for engineering biology, promising leaps forward with beneficial applications. At the same time, these advances raise concerns about biosecurity. Recently, Wittmann et al. described an in silico

A Plan for Global Engagement on AI Standards

April 29, 2025
Author(s)
Jesse Dunietz, Mark Latonero, Kathleen Roberts
This plan has been developed by the Department of Commerce in coordination with the Department of State and agencies across the U.S. Government. It reflects more than 65 comments received in response to a December 2023 Request for Information

2025 NIST GenAI (Pilot) Evaluation Plan for Image Discriminators

March 14, 2025
Author(s)
George Awad, Hariharan Iyer, Seungmin Seo, Peter Fontana, Yooyoung Lee
In this NIST Generative AI (GenAI) program, we invite and encourage participating teams from academia, industry, and other research labs to support research in Generative AI. GenAI is an evaluation series that provides a platform for testing and evaluation

2025 NIST GenAI (Pilot) Evaluation Plan for Image Generators

March 14, 2025
Author(s)
George Awad, Hariharan Iyer, Seungmin Seo, Peter Fontana, Yooyoung Lee
In this NIST Generative AI (GenAI) program, we invite and encourage participating teams from academia, industry, and other research labs to support research in Generative AI. GenAI is an evaluation series that provides a platform for testing and evaluation

Measurement-Based Prediction of mmWave Channel Parameters Using Deep Learning and Point Cloud

August 2, 2024
Author(s)
Anuraag Bodi, Raied Caromi, Jian Wang, Jelena Senic, Camillo Gentile, Hang Mi, Bo Ai, Ruisi He
Millimeter-wave (MmWave) channel characteristics are quite different from sub-6 GHz frequency bands. The major differences include higher path loss and sparser multipath components (MPCs), resulting in more significant time-varying characteristics in

A Plan for Global Engagement on AI Standards

July 26, 2024
Author(s)
Jesse Dunietz, Elham Tabassi, Mark Latonero, Kamie Roberts
Recognizing the importance of technical standards in shaping development and use of Artificial Intelligence (AI), the President's October 2023 Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence (EO 14110)

Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile

July 26, 2024
Author(s)
Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, Kamie Roberts
This document is a cross-sectoral profile of and companion resource for the AI Risk Management Framework (AI RMF 1.0) for Generative AI, pursuant to President Biden's Executive Order (EO) 14110 on Safe, Secure, and Trustworthy Artificial Intelligence. The

On the Evaluation of Machine-Generated Reports

July 14, 2024
Author(s)
James Mayfield, Eugene Yang, Dawn Lawrie, Sean MacAvaney, Paul McNamee, Douglas Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Kayi, Kate Sanders, Marc Mason, Noah Hibbler
Large Language Models (LLMs) have enabled new ways to satisfy information needs. Although great strides have been made in applying them to settings like document ranking and short-form text generation, they still struggle to compose complete, accurate, and

Human-in-the-loop Technical Document Annotation: Developing and Validating a System to Provide Machine-Assistance for Domain-Specific Text Analysis

May 14, 2024
Author(s)
Juan Fung, Zongxia Li, Daniel Stephens, Andrew Mao, Pranav Goel, Emily Walpole, Alden A. Dima, Jordan Boyd-Graber
In this report, we address the following question: to what extent can machine learning assist a human with traditional text analysis, such as content analysis or grounded theory in the social sciences? In practice, such tasks require humans to review and
Was this page helpful?