Prosodic, Intonational, and Voice Quality Correlates of Disfluency
Prosodic, Intonational, and Voice Quality Correlates of Disfluency
批准号:
0414117
负责人:
Mark Hasegawa-Johnson
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-08-15 至 2008-07-31
中文摘要
自发性讲话通常每10-20个单词就会被一个不流利的词打断。虽然人类很容易理解不流利的语音,但计算机语音识别器往往无法将不流利的区域与周围的上下文分开,导致转录失败和意义丧失。该项目通过研究不流利性的感知显著声学关联来开发用于自动语音识别的不流利性识别。在自发性语料库中考察了不流利对音高、能量和声源特征的影响。音调和能量轮廓的近似重复被研究为标记不流利及其后续修复之间的依赖关系的线索。合成分析技术是从语音生成的STEM-ML模型改编而来的,以识别经常被比例差异所掩盖的韵律重复。通过频谱包络的几个声学测量来跟踪语音质量与不流畅的相关性,例如声门化,并执行ROC测试以确定最佳预测因子。通过创建语音语料库的TOBI标准语调标签,考察了不流利性与语调特征、标记口音和短语之间的相关性。在语音识别器的设计中,将不流利的声学和韵律关联与重复语言模型相结合,从而自动转录单词和不流利。识别器集成了多个语言级别的提示,这些提示一起用于识别自发语音中的不流利区域。这项研究将通过识别自然语音中的不流利来推动语音技术的发展。它将为不流利性提供新的统计和声学模型,以及一个可公开访问的带有韵律和不流利性注释的自发语料库。
英文摘要
Spontaneous speech is typically interrupted by one disfluency every 10-20 words. While humans easily comprehend disfluent speech, computer speech recognizers often fail to separate disfluent regions from surrounding context, resulting in failed transcription and loss of meaning. This project develops disfluency recognition for automatic speech recognition by investigating perceptually salient acoustic correlates of disfluency. Effects of disfluency on the pitch, energy and voice source features are examined in the Switchboard corpus of spontaneous speech. The approximate repetition of pitch and energy contours is investigated as a cue marking the dependency between a disfluency and its subsequent repair. Analysis-by-synthesis techniques are adapted from the Stem-ML model of speech generation to recognize prosodic repetition that is often obscured by differences in scaling. Voice quality correlates of disfluency, such as glottalization, are tracked through several acoustic measures of the spectral envelope, with ROC testing performed to determine the best predictors. Correlations between disfluency and intonational features marking accent and phrasing are examined through the creation of a ToBI-standard intonation labeling of the speech corpus. Acoustic and prosodic correlates of disfluency are combined with a repetition language model in the design of a speech recognizer that automatically transcribes both words and disfluencies. The recognizer integrates cues at multiple linguistic levels which together serve to identify regions of disfluency in spontaneous speech. This research will advance speech technology by enabling recognition of disfluency in natural speech. It will contribute new statistical and acoustic models of disfluency and a publicly accessible corpus of spontaneous speech with prosody and disfluency annotation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
-
批准号:2147350
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2022
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
-
批准号:1910319
-
项目类别:Standard Grant
-
资助金额:$25.98万
-
财政年份:2019
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
-
批准号:1550145
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
FODAVA-Partner: Visualizing Audio for Anomaly Detection
-
批准号:0807329
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
-
批准号:0803219
-
项目类别:Standard Grant
-
资助金额:$24.99万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
Audiovisual Distinctive-Feature-Based Recognition of Dysarthric Speech
-
批准号:0534106
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
CAREER: Landmark-Based Speech Recognition in Music and Speech Backgrounds
-
批准号:0132900
-
项目类别:Continuing Grant
-
资助金额:$39.58万
-
财政年份:2002
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
海外基金