课题基金 / 基金详情

RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events

RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
RI Medium:音频二值化 - 全面描述音频事件
批准号:
0803219
负责人:
Mark Hasegawa-Johnson
金额:
$24.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2010-08-31

项目摘要

项目成果

Mark Hasegawa-Johnson的其他基金

相似基金

相关文献

中文摘要
翻译
?知觉突显?视觉心理学家使用这个术语来描述物体吸引观众注意力的能力;例如,已经证明,眼睛运动比不显著的物体更早瞄准显著的物体,并且显著的物体比不显著的物体更快被发现。这项研究的第一个子目标是开发听觉事件知觉显著的自动测量,这里定义为根据幅度、频谱或时间特征(如过零率和周期性)的中心周围对比。这项研究的第二个子目标是使用2007年的伊利诺伊大学Clear评估系统(事件、活动和关系的分类和标签)来测试音频事件检测范式中的显著测量。这项研究的第三个子目标是比较人类标签员观看会议视听记录产生的音频事件转录与不观看任何伴随视频而收听音频的标签者产生的转录;实验假设指出,听觉突显预测仅音频标签比预测视听标签更好。这项研究是计算机视觉和音频信号处理专家之间的合作。如果成功,拟议的方法将有助于为目前安装在许多医院、疗养院、政府大楼和工业现场的视频安全监控系统增加音频通道。
英文摘要
?Perceptual salience? is a term used by psychologists of vision to describe the power of an object to draw viewer attention; for example, it has been demonstrated that eye movements target salient objects sooner than less-salient objects, and that salient objects are detected more quickly than less-salient objects. The first sub-goal of this research is to develop automatic measurements of perceptual salience for auditory events, defined here to be a center-surround contrast in terms of amplitude, spectrum, or temporal features such as zero-crossing rate and periodicity. The second sub-goal of this research is to test salience measurements in an audio event detection paradigm, using the 2007 University of Illinois CLEAR evaluation system (Classification and Labeling of Events, Activities and Relationships). The third sub-goal of this research is to compare audio event transcriptions generated by human labelers viewing an audiovisual record of a meeting vs. transcriptions generated by labelers who listen to the audio without watching any accompanying video; the experimental hypothesis states that auditory salience predicts audio-only labels better than it predicts audiovisual labels. This research is designed as a collaboration between experts in computer vision and audio signal processing. If successful, the proposed methods will help to add an audio channel to the video security monitoring systems currently installed in many hospitals, nursing homes, government buildings and industrial sites.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
FODAVA-Partner: Visualizing Audio for Anomaly Detection
海外基金