Skip to main content
U.S. flag

An official website of the United States government

Official websites use .gov
A .gov website belongs to an official government organization in the United States.

Secure .gov websites use HTTPS
A lock ( ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.

Creating a web-scale video collection for research



Paul D. Over, George M. Awad, Alan Smeaton, Colum Foley, James Lanagan


This paper begins by considering a number of important design questions for a web-scale, widely available, multimedia test collection intended to support long-term scientific evaluation and comparison of content-based video analysis and exploitation systems. Such exploitation systems would include the kinds of functionality already explored within the annual TREC Video Retrieval Evaluation (TRECVid) benchmarking activity such as search, semantic concept detection, and automatic summarization. We then report on our progress in creating such a multimedia collection from publicly available Internet Archive videos with Creative Commons licenses (IACC.1), which we hope will be a useful approximation of a web-scale collection and will support a next generation of benchmarking activities for content-based video operations. We also report on some possibilities for putting this collection to use in multimedia system evaluation.
Proceedings Title
The 1st International Workshop on Web-Scale Multimedia Corpus (WSMC09)
Conference Dates
October 23, 0009-October 23, 2009
Conference Location


benchmarking, evaluation, video retrieval


Over, P. , Awad, G. , Smeaton, A. , Foley, C. and Lanagan, J. (2009), Creating a web-scale video collection for research, The 1st International Workshop on Web-Scale Multimedia Corpus (WSMC09), Beijing, -1, [online], (Accessed June 15, 2024)


If you have any questions about this publication or are having problems accessing it, please contact

Created October 23, 2009, Updated February 19, 2017