RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
批准号:
0803219
负责人:
Mark Hasegawa-Johnson
金额:
$24.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2010-08-31
中文摘要
?知觉显著?是视觉心理学家使用的术语,用来描述一个物体吸引观众注意力的能力;例如,已经证明眼球运动比不太突出的物体更容易瞄准突出的物体,而突出的物体比不太突出的物体更容易被发现。本研究的第一个子目标是开发听觉事件感知显著性的自动测量,这里定义为在振幅,频谱或时间特征(如零交叉率和周期性)方面的中心-环绕对比度。本研究的第二个子目标是测试音频事件检测范式中的显著性测量,使用2007年伊利诺伊大学CLEAR评估系统(事件,活动和关系的分类和标记)。本研究的第三个子目标是比较由观看会议视听记录的人工标记员生成的音频事件转录与不观看任何附带视频的听音频标记员生成的转录;实验假设表明,听觉显著性预测纯音频标签比预测视听标签更好。本研究是计算机视觉和音频信号处理专家之间的合作。如果成功,提议的方法将有助于为目前安装在许多医院、养老院、政府大楼和工业场所的视频安全监控系统增加一个音频通道。
英文摘要
?Perceptual salience? is a term used by psychologists of vision to describe the power of an object to draw viewer attention; for example, it has been demonstrated that eye movements target salient objects sooner than less-salient objects, and that salient objects are detected more quickly than less-salient objects. The first sub-goal of this research is to develop automatic measurements of perceptual salience for auditory events, defined here to be a center-surround contrast in terms of amplitude, spectrum, or temporal features such as zero-crossing rate and periodicity. The second sub-goal of this research is to test salience measurements in an audio event detection paradigm, using the 2007 University of Illinois CLEAR evaluation system (Classification and Labeling of Events, Activities and Relationships). The third sub-goal of this research is to compare audio event transcriptions generated by human labelers viewing an audiovisual record of a meeting vs. transcriptions generated by labelers who listen to the audio without watching any accompanying video; the experimental hypothesis states that auditory salience predicts audio-only labels better than it predicts audiovisual labels. This research is designed as a collaboration between experts in computer vision and audio signal processing. If successful, the proposed methods will help to add an audio channel to the video security monitoring systems currently installed in many hospitals, nursing homes, government buildings and industrial sites.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
-
批准号:2147350
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2022
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
-
批准号:1910319
-
项目类别:Standard Grant
-
资助金额:$25.98万
-
财政年份:2019
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
-
批准号:1550145
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
FODAVA-Partner: Visualizing Audio for Anomaly Detection
-
批准号:0807329
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
Audiovisual Distinctive-Feature-Based Recognition of Dysarthric Speech
-
批准号:0534106
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
Prosodic, Intonational, and Voice Quality Correlates of Disfluency
-
批准号:0414117
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
CAREER: Landmark-Based Speech Recognition in Music and Speech Backgrounds
-
批准号:0132900
-
项目类别:Continuing Grant
-
资助金额:$39.58万
-
财政年份:2002
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
海外基金