CI-P: Planning for AudioNet: A New Community Infrastructure for Audio Annotations for Acoustic Event Identification
CI-P: Planning for AudioNet: A New Community Infrastructure for Audio Annotations for Acoustic Event Identification
批准号:
1629990
负责人:
Gerald Friedland
金额:
$10.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-07-01 至 2018-12-31
中文摘要
这项工作为AudioNet奠定了基础,AudioNet是一个公共领域的音频标签语料库,用于开放访问的YFCC100M数据集中的近80万个视频。在视频数据的自动分析中,音频信息为视觉信息提供了重要的补充,使系统能够检测到仅从视觉流可能无法清楚识别的情况。然而,到目前为止,还没有真正大规模的标记音频数据集,需要作为输入来构建灵活、准确的分析系统。创建这样一个大规模的语料库将推动更多的研究人员和计算机科学专业的学生开发更好的多媒体算法,从而转化为对公众日常生活的影响。社交媒体视频越来越多地用于科学研究,因为它们为观察和模拟社会科学、经济学、气象学和医学中的许多现象提供了机会。因此,内容分析的新功能将影响许多科学领域。此外,音频分析还可以用于实时安全监控,以及自动驾驶汽车和家用机器人等机器人应用,以帮助和监控老年人。AudioNet是一个多机构合作的多媒体共享计划的一部分,该计划正在围绕知识共享协议许可的照片和视频的YFCC100M数据集开发各种资源。AudioNet正在对YFCC100M视频中的音轨进行注释,重点关注音频概念。音频概念可以被视为声音“对象”:具体的、可定位的声音单位,如“人群欢呼”或“火警”。该方法将在ImageNet上建模,ImageNet是一个使用WordNet语法集层次结构(同义词组)标记和组织的图像数据集;ImageNet在图像处理方面取得了重大进展。然而,虽然ImageNet主要关注实体(名词同义词集),但音频数据本质上是暂时的。因此,AudioNet的标签集将侧重于事件和操作,尽管使用类似WordNet的语义资源进行组织。
英文摘要
This effort lays the groundwork for AudioNet, a public-domain corpus of audio labels for the nearly 800,000 videos in the open-access YFCC100M dataset. Audio information provides an important complement to visual information in the automatic analysis of video data, allowing systems to detect situations that may not be clearly identifiable from the visual stream alone. However, there are as yet no truly large-scale labeled audio datasets of the kind needed as input to build flexible, accurate analysis systems. Creating such a large-scale corpus will serve as an impetus for better multimedia algorithms to be developed by more researchers and computer science students, translating into an impact on the everyday life of the public at large. Social media videos are increasingly used for scientific research, as they provide an opportunity to observe and model many phenomena in the social sciences, economics, meteorology, and medicine. New capabilities for content analysis will therefore impact many scientific fields. In addition, audio analysis could be used in real-time security surveillance and in robotics applications like autonomous vehicles and household robots to aid and monitor the elderly.AudioNet is part of a multi-institution collaboration, the Multimedia Commons initiative, which is developing a variety of resources around the YFCC100M dataset of Creative Commons-licensed photos and videos. AudioNet is annotating the audio tracks from the YFCC100M videos, focusing on audio concepts. Audio concepts can be thought of as acoustic "objects": concrete, localizable units of sound like "crowd cheering" or "fire alarm". The approach will be modeled on ImageNet, an image dataset labeled and organized using the WordNet hierarchy of synsets (groups of synonyms); ImageNet has enabled major enabled advances in image processing. However, while ImageNet focuses largely on entities (noun synsets), audio data is inherently temporal. The label set for AudioNet will therefore focus on events and actions, though similarly organized using semantic resources like WordNet.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EDU: Teachers' Resources for Online Privacy Education (TROPE)
-
批准号:1419319
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2014
-
负责人:Gerald Friedland
-
依托单位:
I-Corps: Commercializing the Integration of Human and Artificial Intelligence for Large Scale Multimedia Analysis
-
批准号:1339552
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2013
-
负责人:Gerald Friedland
-
依托单位:
BIGDATA: Small: DCM: DA: Collaborative Research: SMASH -- Scalable Multimedia content AnalysiS in a High-level language
-
批准号:1251276
-
项目类别:Standard Grant
-
资助金额:$40.4万
-
财政年份:2013
-
负责人:Gerald Friedland
-
依托单位:
EAGER: Collecting Training Videos for Location Estimation with Mechanical Turk
-
批准号:1138599
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2011
-
负责人:Gerald Friedland
-
依托单位:
海外基金