Prosodic, Intonational, and Voice Quality Correlates of Disfluency
Prosodic, Intonational, and Voice Quality Correlates of Disfluency
批准号:
0414117
负责人:
Mark Hasegawa-Johnson
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-08-15 至 2008-07-31
中文摘要
自发言语通常每10-20个单词就会被一次不流利打断。虽然人类很容易理解不流利的语音,但计算机语音识别器通常无法将不流利的区域与周围的上下文分开,导致转录失败和意义丧失。 本计画借由研究不流利的听觉相关性,发展自动语音辨识中的不流利辨识。 在Switchboard自发语音语料库中考察了不流利性对音高、能量和声源特征的影响。 音高和能量轮廓的近似重复作为一个线索标记之间的依赖性不流利和随后的修复进行了研究。 合成分析技术适用于语音生成的Stem-ML模型,以识别通常被缩放差异所掩盖的韵律重复。 语音质量相关的不流利,如声门化,通过几个声学测量的频谱包络跟踪,与ROC测试进行,以确定最佳的预测。 通过建立一个符合ToBI标准的语音语调标注语料库,研究了语调特征、重音和措辞与不流利性之间的相关性。 不流利的声学和韵律相关性与重复语言模型相结合,设计了一个自动转录单词和不流利的语音识别器。 识别器整合了多个语言层次的线索,这些线索一起用于识别自发言语中的不流利区域。这项研究将通过识别自然语音中的不流利来推进语音技术。 它将提供新的统计和不流利的声学模型和一个公开访问的语料库的自发语音与韵律和不流利的注释。
英文摘要
Spontaneous speech is typically interrupted by one disfluency every 10-20 words. While humans easily comprehend disfluent speech, computer speech recognizers often fail to separate disfluent regions from surrounding context, resulting in failed transcription and loss of meaning. This project develops disfluency recognition for automatic speech recognition by investigating perceptually salient acoustic correlates of disfluency. Effects of disfluency on the pitch, energy and voice source features are examined in the Switchboard corpus of spontaneous speech. The approximate repetition of pitch and energy contours is investigated as a cue marking the dependency between a disfluency and its subsequent repair. Analysis-by-synthesis techniques are adapted from the Stem-ML model of speech generation to recognize prosodic repetition that is often obscured by differences in scaling. Voice quality correlates of disfluency, such as glottalization, are tracked through several acoustic measures of the spectral envelope, with ROC testing performed to determine the best predictors. Correlations between disfluency and intonational features marking accent and phrasing are examined through the creation of a ToBI-standard intonation labeling of the speech corpus. Acoustic and prosodic correlates of disfluency are combined with a repetition language model in the design of a speech recognizer that automatically transcribes both words and disfluencies. The recognizer integrates cues at multiple linguistic levels which together serve to identify regions of disfluency in spontaneous speech. This research will advance speech technology by enabling recognition of disfluency in natural speech. It will contribute new statistical and acoustic models of disfluency and a publicly accessible corpus of spontaneous speech with prosody and disfluency annotation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
-
批准号:2147350
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2022
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
-
批准号:1910319
-
项目类别:Standard Grant
-
资助金额:$25.98万
-
财政年份:2019
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
-
批准号:1550145
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
FODAVA-Partner: Visualizing Audio for Anomaly Detection
-
批准号:0807329
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
-
批准号:0803219
-
项目类别:Standard Grant
-
资助金额:$24.99万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
Audiovisual Distinctive-Feature-Based Recognition of Dysarthric Speech
-
批准号:0534106
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
CAREER: Landmark-Based Speech Recognition in Music and Speech Backgrounds
-
批准号:0132900
-
项目类别:Continuing Grant
-
资助金额:$39.58万
-
财政年份:2002
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
海外基金