Some of our key past projects, ordered by the most recent completion date, include:
- DARPA Computational Cultural Understanding (CCU) [2022 - 2025]: DARPA's CCU program set out to build human language technologies that help monolingual operators handle cross-cultural dialogue effectively. NIST evaluated system performance across two technical areas: Sociocultural Analysis, covering component technologies for identifying sociocultural norms, recognizing emotions across cultures, and spotting significant shifts in norms and emotional expression; and Cross-Cultural Dialogue Assistance, which developed a framework for a dialogue assistant supporting monolingual operators in cross-cultural settings.
- IARPA Machine Translation for English Retrieval of Information in Any Language (MATERIAL) [2018 - 2021]: MATERIAL aimed to build systems that could take domain-contextualized English queries, locate relevant speech or text in low-resource languages, and generate English summaries to help analysts triage information efficiently. Systems had to work with limited bitext data and no domain adaptation data, yet still generalize to new domains and genres, producing static text summaries that let users judge relevance at a glance. NIST designed and conducted the program's performance evaluations.
- Open Speech Analytic Technologies (OpenSAT) [2017 - 2020]: OpenSAT brought together researchers tackling speech analytics in difficult acoustic conditions, using large-scale, objective evaluations to advance the field. Its core focus areas were speech activity detection (SAD), automatic speech recognition (ASR), and keyword search (KWS).
- Activities in Extended Videos (ActEv) [2018 - 2019]: ActEV benchmarked robust automatic activity detection algorithms for multi-camera streaming video environments, with applications spanning both forensic analysis and real-time alerting.
- DARPA Low Resource Languages for Emergent Incidents (LORELEI) and Low Resource Human Language Technologies (LoReHLT) [2016-2019]: LORELEI sought to dramatically advance computational linguistics and human language technology to enable rapid, low-cost development of capabilities for low-resource languages. LoReHLT served as LORELEI's public counterpart, opening participation in the evaluation to researchers outside the core program.
- Media Forensics Challenge [2017 - 2020]: This challenge promoted technologies that could automatically assess the integrity of images and video, determining whether media had been authentically captured or manipulated.
- Multimedia Event Detection (MED) [2010 - 2017]: MED aimed to assemble core detection technologies into systems capable of searching multimedia recordings for user-defined events based on pre-computed metadata, with that metadata designed to be general enough for reuse across future, unforeseen ad-hoc events.
- Surveillance Event Detection (SED) [2008 - 2017]: SED promoted the development of technologies for detecting activities occurring within surveillance video.
- DARPA Broad Operational Language Translation (BOLT) [2011-2015]: The DARPA BOLT program aimed at enabling English-speaking persons to communicate with non-English-speaking populations and identify important information in foreign-language sources -- specifically by: 1) allowing English-speakers to understand foreign-language sources of all genres, including chat, messaging and informal conversation; 2) providing English-speakers the ability to quickly identify targeted information in foreign-language sources using natural-language queries; and 3) enabling multi-turn communication in text and speech with non-English speakers. There was some emphasis on delivering these capabilities free from domain or genre limitations. NIST organized and implemented BOLT's evaluations of speech-to-text and text-to-text MT technology as well as the evaluation of end-to-end MT systems enabling live speech communication between two speakers of different languages.
- Open Machine Translation (OpenMT) [2000 - 2015]: A biannual NIST evaluation series, OpenMT focused on core text-to-text MT tasks designed so that findings could transfer to other machine translation applications.
- DARPA Multilingual Automatic Document Classification Analysis and Translation (MADCAT) [2008 - 2013] and OpenHaRT [2011-2013]: MADCAT aimed to automatically convert foreign-language text images, especially handwritten Arabic, into English transcripts for use by monolingual English speakers. NIST organized and ran yearly evaluations using both controlled data sets and real-life data, later extending this evaluation model to the public through OpenHaRT.
- DARPA Global Autonomous Language Exploitation (GALE) [2006-2011]: GALE aimed to absorb, translate, analyze, and interpret large volumes of speech and text across multiple languages, particularly Arabic and Mandarin Chinese, producing tailored distillations to help military personnel and analysts overcome language barriers. NIST served as GALE's independent evaluator, running yearly MT evaluations built around an "edit-distance" metric measuring how many edits were needed to make output fluent and accurate, and separately evaluating distillation technology by having annotators judge response relevance and count "nuggets" of atomic information, an approach adapted from NIST's complex question-answering methodology.
- AVSS Multiple Camera Single Person Tracking (MCSPT) Challenge [2009 - 2010]: MCSPT provided a common evaluation task focused on tracking a specified person across a video sensor field using a small set of in situ exemplar images.
- DARPA Spoken Language Communication and Translation System for Tactical Use (TRANSTAC) [2006-2010]: TRANSTAC aimed to develop and field speech-to-speech machine translation technology enabling two-way spoken communication between English-speaking U.S. Soldiers and Marines and civilian populations speaking other languages. NIST evaluated the TRANSTAC systems as a whole, as well as their individual speech recognition, machine translation, and text-to-speech components.
- Video Surveillance Technologies for Retail Security (VISITORS) [2010]: VISITORS advanced predictive analysis technologies for detecting people engaged in suspicious activity in surveillance video, with a focus on the retail domain.
- Classification of Events, Activities and Relationships (CLEAR): CLEAR was a multi-national evaluation series joining researchers from the US ARDA VACE Program and the EU's Computers in the Human Interaction Loop Program to advance detection and tracking of people, faces, vehicles, and related targets, along with acoustic event detection.
- DARPA Data-Driven Discoveries of Models (D3M): NIST supported DARPA's DARPA D3M program as part of its Test and Evaluation Government Team; D3M aimed to let automated systems, paired with subject matter experts, model and solve complex machine learning problems.
- DARPA XDATA: NIST supported XDATA's development of tools for handling the computational challenges of analyzing large, incomplete data sets by evaluating the software tools produced, providing standardized reports on analytic accuracy and system benchmarking. Many resulting tools were released through the DARPA Open Catalog.
- Data Science Evaluation Series (DSE): DSE offered a cross-disciplinary framework for evaluating data analytic algorithms across the entire analytic pipeline, giving domain-specific algorithms a generalized setting in which to be adopted and assessed.
- Evaluation Management System (EMS): The EMS project provided evaluation infrastructure for the DSE that allowed for isolated environments to evaluate submissions, advanced benchmarking analytics, and a platform to run systems that use a variety of architectures and distributed programming frameworks (such as Hadoop and Spark). The EMS integrated hardware and software components for easy deployment and reconfiguration of computational needs and enabled integration of compute-and data-intensive problems within a controlled private cloud. This design allowed for test and evaluation of different compute paradigms as well as facilitated the integration of hardware acceleration components in order to best assess how a given evaluation could be run.
- Metrics Matter (MetricsMaTr): MetricsMaTr focused on developing automated measurement techniques for MT technology, aimed at providing insight into translation quality.
- Machine Foreign Language Translation System (MFLTS): Sponsored by the US Army's MFLTS program, this project had NIST chair the Metrics-IPT working group tasked with developing a new metric grounded in the Interagency Language Roundtable (ILR) rating system.
- NIST Data Science Symposium: This symposium, drawing over 700 attendees, provided a forum for discussing measurement science for data analytics, with NIST helping organize and coordinate the event.
- Open Speech Activity Detection (OpenSAD): OpenSAD benchmarked speech activity detection systems across varied audio data, serving as an open counterpart to the DARPA RATS SAD evaluations and welcoming all interested participants.
- Open Key Word Search (OpenKWS): An annual evaluation of keyword search technologies in a new language each year, OpenKWS grew out of the 2006 Spoken Term Detection evaluation.
- Rich Transcription: The Rich Transcription evaluation series promoted and measured advances across several automatic speech recognition technologies, aiming to produce transcriptions that were more readable for humans and more useful for machines.
- Video Analysis and Content Extraction (VACE): Established under the US Advanced Research and Development Activity (ARDA), VACE developed novel algorithms for automatic video content extraction, multi-modal fusion, and event understanding, advancing automated detection and tracking of moving objects including faces, hands, people, vehicles, and text.