CI-P: Planning for AudioNet: A New Community Infrastructure for Audio Annotations for Acoustic Event Identification
CI-P: Planning for AudioNet: A New Community Infrastructure for Audio Annotations for Acoustic Event Identification
批准号:
1629990
负责人:
Gerald Friedland
金额:
$10.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-07-01 至 2018-12-31
中文摘要
这一努力为AudioNet奠定了基础,AudioNet是一个公共领域的音频标签语料库,涵盖开放访问的YFCC100M数据集中的近80万个视频。在视频数据的自动分析中,音频信息为视觉信息提供了重要的补充,使系统能够检测到仅从视频流中可能无法清楚识别的情况。然而,到目前为止,还没有真正大规模的标记音频数据集,作为建立灵活、准确的分析系统所需的输入。创建如此大规模的语料库将推动更多的研究人员和计算机科学专业的学生开发更好的多媒体算法,转化为对广大公众日常生活的影响。社交媒体视频越来越多地被用于科学研究,因为它们提供了观察社会科学、经济学、气象学和医学中的许多现象并建立模型的机会。因此,内容分析的新能力将影响许多科学领域。此外,音频分析还可以用于实时安全监控,以及自动驾驶车辆和家用机器人等机器人应用程序,以帮助和监控老年人。AudioNet是多机构合作的一部分,即多媒体共享计划,该计划正在围绕YFCC100M授权照片和视频的YFCC100M数据集开发各种资源。AudioNet正在为YFCC100M视频中的音轨添加注释,重点放在音频概念上。音频概念可以被认为是声学“对象”:具体的、可本地化的声音单位,如“人群欢呼”或“火警”。该方法将以ImageNet为模型,这是一种使用同义词集(同义词组)的WordNet层次结构标记和组织的图像数据集;ImageNet在图像处理方面取得了重大进展。然而,虽然ImageNet主要关注实体(名词同义词集),但音频数据本质上是临时的。因此,AudioNet的标签集将专注于事件和动作,尽管类似地使用WordNet等语义资源进行组织。
英文摘要
This effort lays the groundwork for AudioNet, a public-domain corpus of audio labels for the nearly 800,000 videos in the open-access YFCC100M dataset. Audio information provides an important complement to visual information in the automatic analysis of video data, allowing systems to detect situations that may not be clearly identifiable from the visual stream alone. However, there are as yet no truly large-scale labeled audio datasets of the kind needed as input to build flexible, accurate analysis systems. Creating such a large-scale corpus will serve as an impetus for better multimedia algorithms to be developed by more researchers and computer science students, translating into an impact on the everyday life of the public at large. Social media videos are increasingly used for scientific research, as they provide an opportunity to observe and model many phenomena in the social sciences, economics, meteorology, and medicine. New capabilities for content analysis will therefore impact many scientific fields. In addition, audio analysis could be used in real-time security surveillance and in robotics applications like autonomous vehicles and household robots to aid and monitor the elderly.AudioNet is part of a multi-institution collaboration, the Multimedia Commons initiative, which is developing a variety of resources around the YFCC100M dataset of Creative Commons-licensed photos and videos. AudioNet is annotating the audio tracks from the YFCC100M videos, focusing on audio concepts. Audio concepts can be thought of as acoustic "objects": concrete, localizable units of sound like "crowd cheering" or "fire alarm". The approach will be modeled on ImageNet, an image dataset labeled and organized using the WordNet hierarchy of synsets (groups of synonyms); ImageNet has enabled major enabled advances in image processing. However, while ImageNet focuses largely on entities (noun synsets), audio data is inherently temporal. The label set for AudioNet will therefore focus on events and actions, though similarly organized using semantic resources like WordNet.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EDU: Teachers' Resources for Online Privacy Education (TROPE)
-
批准号:1419319
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2014
-
负责人:Gerald Friedland
-
依托单位:
I-Corps: Commercializing the Integration of Human and Artificial Intelligence for Large Scale Multimedia Analysis
-
批准号:1339552
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2013
-
负责人:Gerald Friedland
-
依托单位:
BIGDATA: Small: DCM: DA: Collaborative Research: SMASH -- Scalable Multimedia content AnalysiS in a High-level language
-
批准号:1251276
-
项目类别:Standard Grant
-
资助金额:$40.4万
-
财政年份:2013
-
负责人:Gerald Friedland
-
依托单位:
EAGER: Collecting Training Videos for Location Estimation with Mechanical Turk
-
批准号:1138599
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2011
-
负责人:Gerald Friedland
-
依托单位:
海外基金