课题基金 / 基金详情

SGER: Incorporating Higher-Level Information into Dynamic Pronounciation Modeling for ASR

SGER: Incorporating Higher-Level Information into Dynamic Pronounciation Modeling for ASR
SGER:将高级信息纳入 ASR 动态发音建模
批准号:
9713346
负责人:
Nelson Morgan
金额:
$3.81万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-10-01 至 1998-09-30

项目摘要

项目成果

Nelson Morgan的其他基金

相似基金

相关文献

中文摘要
翻译
在大词汇量的自发言语中,单词发音的变异性远高于朗读言语情景。在1996年的大词汇量会话语音识别夏季研讨会(WS96)上,基于语音数据的机器派生描述,开发了用于自动语音识别(ASR)系统的这种可变性的模型。这项研究的继续工作侧重于研究连续语音中发音变化与ASR发音模型中通常没有的更高级别信息之间的相关性。这个模型中的一个重要因素是语音率,它已经被证明是阅读和自然语音语料库中词错误率的一个很好的预测因子。研究了重音化(单词顺序发音时音节边界的移动)和词频对单词发音的影响。本项目的目标是提高语音识别模型的可预测性,特别是降低自然语音和会话语音的识别错误。这些技术将在Switchboard语料库上进行评估。
英文摘要
In large-vocabulary spontaneous speech, the variability of the pronunciations of words is much higher than in read speech situations. At the 1996 Summer Workshop on Large Vocabulary Conversational Speech Recognition (WS96), a model for this variability to be used in Automatic Speech Recognition (ASR) systems was developed based on machine- derived descriptions of speech data. The continuation of this work in this grant focuses on studying the correlation of variation in pronunciations in continuous speech and higher-level information not usually brought to bear in an ASR pronunciation model. One important element in this model is the rate of speech, which has been shown to be a good predictor of word error rate on both read and spontaneous speech corpora. Investigations into the effects of resyllabification (movement of syllable boundaries when words are spoken in sequence) and word frequency on word pronunciations are also undertaken. The goal of this project is to improve the predictability of variation for speech recognition models, in particular for the reduction of recognition error for spontaneous and conversational speech. The techniques will be evaluated on the Switchboard corpus.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Collaborative Research: Towards Modeling Source Separation from Measured Cortical Responses
EAGER: Collaborative Research: Towards Modeling Human Speech Confusions in Noise
International: An Analysis of Speaker Diarization Systems Errors
CI-P: Towards a Consensus Representation for Understanding Structure of Multiparty Conversations
海外基金