课题基金 / 基金详情

SGER: Incorporating Higher-Level Information into Dynamic Pronounciation Modeling for ASR

SGER: Incorporating Higher-Level Information into Dynamic Pronounciation Modeling for ASR
SGER:将高级信息纳入 ASR 动态发音建模
批准号:
9713346
负责人:
Nelson Morgan
金额:
$3.81万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-10-01 至 1998-09-30

项目摘要

项目成果

Nelson Morgan的其他基金

相似基金

相关文献

中文摘要
翻译
在大词汇量的自发言语中,单词发音的变异性要比在朗读的情况下高得多。 在1996年的大词汇量会话语音识别(WS96)夏季研讨会上, 在自动语音识别(ASR)系统中使用的这种可变性是基于语音数据的机器导出的描述而开发的。 这项工作的继续在这个补助金的重点是研究在连续语音和更高层次的信息,通常不承担在ASR发音模型中的发音变化的相关性。 这个模型中的一个重要元素是语音速率,它已被证明是一个很好的预测词错误率在两个阅读 和自发语音语料库。 调查的影响resyllabification(移动音节边界时,单词顺序发言)和单词的发音频率也进行。 该项目的目标是提高语音识别模型变化的可预测性,特别是减少自发和会话语音的识别错误。 这些技术将在Switchboard语料库上进行评估。
英文摘要
In large-vocabulary spontaneous speech, the variability of the pronunciations of words is much higher than in read speech situations. At the 1996 Summer Workshop on Large Vocabulary Conversational Speech Recognition (WS96), a model for this variability to be used in Automatic Speech Recognition (ASR) systems was developed based on machine- derived descriptions of speech data. The continuation of this work in this grant focuses on studying the correlation of variation in pronunciations in continuous speech and higher-level information not usually brought to bear in an ASR pronunciation model. One important element in this model is the rate of speech, which has been shown to be a good predictor of word error rate on both read and spontaneous speech corpora. Investigations into the effects of resyllabification (movement of syllable boundaries when words are spoken in sequence) and word frequency on word pronunciations are also undertaken. The goal of this project is to improve the predictability of variation for speech recognition models, in particular for the reduction of recognition error for spontaneous and conversational speech. The techniques will be evaluated on the Switchboard corpus.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Collaborative Research: Towards Modeling Source Separation from Measured Cortical Responses
EAGER: Collaborative Research: Towards Modeling Human Speech Confusions in Noise
International: An Analysis of Speaker Diarization Systems Errors
CI-P: Towards a Consensus Representation for Understanding Structure of Multiparty Conversations
海外基金