CI-P: Planning for AudioNet: A New Community Infrastructure for Audio Annotations for Acoustic Event Identification
CI-P: Planning for AudioNet: A New Community Infrastructure for Audio Annotations for Acoustic Event Identification
批准号:
1629990
负责人:
Gerald Friedland
金额:
$10.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-07-01 至 2018-12-31
中文摘要
这项工作为AudioNet奠定了基础,AudioNet是一个公共领域的音频标签语料库,用于开放访问的YFCC 100M数据集中的近80万个视频。在视频数据的自动分析中,音频信息为视觉信息提供了重要的补充,使系统能够检测可能无法单独从视觉流中清楚识别的情况。然而,目前还没有真正大规模的标记音频数据集作为构建灵活,准确的分析系统所需的输入。创建这样一个大规模的语料库将推动更多的研究人员和计算机科学专业的学生开发更好的多媒体算法,从而对公众的日常生活产生影响。社交媒体视频越来越多地用于科学研究,因为它们提供了观察和建模社会科学,经济学,气象学和医学中许多现象的机会。因此,内容分析的新能力将影响许多科学领域。此外,音频分析还可用于实时安全监控和机器人应用,如自动驾驶汽车和家用机器人,以帮助和监控老年人。AudioNet是多机构合作的多媒体共享计划的一部分,该计划正在围绕YFCC 100M数据集开发各种资源,这些数据集包含知识共享许可的照片和视频。AudioNet正在注释YFCC 100M视频的音轨,重点是音频概念。音频概念可以被认为是声学“对象”:具体的、可定位的声音单位,如“人群欢呼”或“火警”。该方法将在ImageNet上建模,ImageNet是一个使用同义词集(同义词组)的WordNet层次结构标记和组织的图像数据集; ImageNet已经实现了图像处理的重大进步。然而,虽然ImageNet主要关注实体(名词同义词),但音频数据本质上是时间性的。因此,AudioNet的标签集将专注于事件和动作,尽管类似地使用WordNet等语义资源进行组织。
英文摘要
This effort lays the groundwork for AudioNet, a public-domain corpus of audio labels for the nearly 800,000 videos in the open-access YFCC100M dataset. Audio information provides an important complement to visual information in the automatic analysis of video data, allowing systems to detect situations that may not be clearly identifiable from the visual stream alone. However, there are as yet no truly large-scale labeled audio datasets of the kind needed as input to build flexible, accurate analysis systems. Creating such a large-scale corpus will serve as an impetus for better multimedia algorithms to be developed by more researchers and computer science students, translating into an impact on the everyday life of the public at large. Social media videos are increasingly used for scientific research, as they provide an opportunity to observe and model many phenomena in the social sciences, economics, meteorology, and medicine. New capabilities for content analysis will therefore impact many scientific fields. In addition, audio analysis could be used in real-time security surveillance and in robotics applications like autonomous vehicles and household robots to aid and monitor the elderly.AudioNet is part of a multi-institution collaboration, the Multimedia Commons initiative, which is developing a variety of resources around the YFCC100M dataset of Creative Commons-licensed photos and videos. AudioNet is annotating the audio tracks from the YFCC100M videos, focusing on audio concepts. Audio concepts can be thought of as acoustic "objects": concrete, localizable units of sound like "crowd cheering" or "fire alarm". The approach will be modeled on ImageNet, an image dataset labeled and organized using the WordNet hierarchy of synsets (groups of synonyms); ImageNet has enabled major enabled advances in image processing. However, while ImageNet focuses largely on entities (noun synsets), audio data is inherently temporal. The label set for AudioNet will therefore focus on events and actions, though similarly organized using semantic resources like WordNet.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EDU: Teachers' Resources for Online Privacy Education (TROPE)
-
批准号:1419319
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2014
-
负责人:Gerald Friedland
-
依托单位:
I-Corps: Commercializing the Integration of Human and Artificial Intelligence for Large Scale Multimedia Analysis
-
批准号:1339552
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2013
-
负责人:Gerald Friedland
-
依托单位:
BIGDATA: Small: DCM: DA: Collaborative Research: SMASH -- Scalable Multimedia content AnalysiS in a High-level language
-
批准号:1251276
-
项目类别:Standard Grant
-
资助金额:$40.4万
-
财政年份:2013
-
负责人:Gerald Friedland
-
依托单位:
EAGER: Collecting Training Videos for Location Estimation with Mechanical Turk
-
批准号:1138599
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2011
-
负责人:Gerald Friedland
-
依托单位:
海外基金