Prosodic, Intonational, and Voice Quality Correlates of Disfluency
Prosodic, Intonational, and Voice Quality Correlates of Disfluency
批准号:
0414117
负责人:
Mark Hasegawa-Johnson
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-08-15 至 2008-07-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Spontaneous speech is typically interrupted by one disfluency every 10-20 words. While humans easily comprehend disfluent speech, computer speech recognizers often fail to separate disfluent regions from surrounding context, resulting in failed transcription and loss of meaning. This project develops disfluency recognition for automatic speech recognition by investigating perceptually salient acoustic correlates of disfluency. Effects of disfluency on the pitch, energy and voice source features are examined in the Switchboard corpus of spontaneous speech. The approximate repetition of pitch and energy contours is investigated as a cue marking the dependency between a disfluency and its subsequent repair. Analysis-by-synthesis techniques are adapted from the Stem-ML model of speech generation to recognize prosodic repetition that is often obscured by differences in scaling. Voice quality correlates of disfluency, such as glottalization, are tracked through several acoustic measures of the spectral envelope, with ROC testing performed to determine the best predictors. Correlations between disfluency and intonational features marking accent and phrasing are examined through the creation of a ToBI-standard intonation labeling of the speech corpus. Acoustic and prosodic correlates of disfluency are combined with a repetition language model in the design of a speech recognizer that automatically transcribes both words and disfluencies. The recognizer integrates cues at multiple linguistic levels which together serve to identify regions of disfluency in spontaneous speech. This research will advance speech technology by enabling recognition of disfluency in natural speech. It will contribute new statistical and acoustic models of disfluency and a publicly accessible corpus of spontaneous speech with prosody and disfluency annotation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
-
批准号:2147350
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2022
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
-
批准号:1910319
-
项目类别:Standard Grant
-
资助金额:$25.98万
-
财政年份:2019
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
-
批准号:1550145
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
FODAVA-Partner: Visualizing Audio for Anomaly Detection
-
批准号:0807329
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
-
批准号:0803219
-
项目类别:Standard Grant
-
资助金额:$24.99万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
Audiovisual Distinctive-Feature-Based Recognition of Dysarthric Speech
-
批准号:0534106
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
CAREER: Landmark-Based Speech Recognition in Music and Speech Backgrounds
-
批准号:0132900
-
项目类别:Continuing Grant
-
资助金额:$39.58万
-
财政年份:2002
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
海外基金