课题基金 / 基金详情

RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events

RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
RI Medium:音频二值化 - 全面描述音频事件
批准号:
0803219
负责人:
Mark Hasegawa-Johnson
金额:
$24.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2010-08-31

项目摘要

项目成果

Mark Hasegawa-Johnson的其他基金

相似基金

相关文献

中文摘要
翻译
?知觉显著?是视觉心理学家使用的术语,用来描述一个物体吸引观众注意力的能力;例如,已经证明眼球运动比不太突出的物体更容易瞄准突出的物体,而突出的物体比不太突出的物体更容易被发现。本研究的第一个子目标是开发听觉事件感知显著性的自动测量,这里定义为在振幅,频谱或时间特征(如零交叉率和周期性)方面的中心-环绕对比度。本研究的第二个子目标是测试音频事件检测范式中的显著性测量,使用2007年伊利诺伊大学CLEAR评估系统(事件,活动和关系的分类和标记)。本研究的第三个子目标是比较由观看会议视听记录的人工标记员生成的音频事件转录与不观看任何附带视频的听音频标记员生成的转录;实验假设表明,听觉显著性预测纯音频标签比预测视听标签更好。本研究是计算机视觉和音频信号处理专家之间的合作。如果成功,提议的方法将有助于为目前安装在许多医院、养老院、政府大楼和工业场所的视频安全监控系统增加一个音频通道。
英文摘要
?Perceptual salience? is a term used by psychologists of vision to describe the power of an object to draw viewer attention; for example, it has been demonstrated that eye movements target salient objects sooner than less-salient objects, and that salient objects are detected more quickly than less-salient objects. The first sub-goal of this research is to develop automatic measurements of perceptual salience for auditory events, defined here to be a center-surround contrast in terms of amplitude, spectrum, or temporal features such as zero-crossing rate and periodicity. The second sub-goal of this research is to test salience measurements in an audio event detection paradigm, using the 2007 University of Illinois CLEAR evaluation system (Classification and Labeling of Events, Activities and Relationships). The third sub-goal of this research is to compare audio event transcriptions generated by human labelers viewing an audiovisual record of a meeting vs. transcriptions generated by labelers who listen to the audio without watching any accompanying video; the experimental hypothesis states that auditory salience predicts audio-only labels better than it predicts audiovisual labels. This research is designed as a collaboration between experts in computer vision and audio signal processing. If successful, the proposed methods will help to add an audio channel to the video security monitoring systems currently installed in many hospitals, nursing homes, government buildings and industrial sites.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
FODAVA-Partner: Visualizing Audio for Anomaly Detection
海外基金