课题基金 / 基金详情

Automatic Speech Recognition Based on Syllable-length Acoustic Models

Automatic Speech Recognition Based on Syllable-length Acoustic Models
基于音节长度声学模型的自动语音识别
批准号:
9712579
负责人:
Nelson Morgan
金额:
$78.45万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-09-15 至 2000-08-31

项目摘要

项目成果

Nelson Morgan的其他基金

相似基金

相关文献

中文摘要
翻译
与传统电话或子电话相比,具有更长的时间基础的单元可能为为识别自然口语话语而定制的ASR系统提供更好的结构基础。在这个项目中,音节长度声学模型正在被探索用于识别会话语音。该项目需要:定义由跨越约250毫秒语音间隔的能量轨迹衍生的声学特征;音节长度区域的统计建模;开发一种解码方案,旨在将这些声学模型的输出与更传统的电话和子电话长度模型相结合;并将音节、电话和子电话特征嵌入到语言的多层表示中,以便在广泛的声学环境和说话条件下进行鲁棒识别。一个完整的ASR系统正在开发中,该系统将纳入本研究的结果,并将对流利的语言进行评估。识别系统的成功结果有可能改善必须处理自发话语解码的实际ASR系统。此外,在声学、统计和词汇层面上对长时结构的分析可以提高我们对会话语音结构的基本认识。
英文摘要
Units with a longer time basis than the traditional phones or sub- phones may provide a better structural basis for ASR systems customized for recognition of naturally spoken discourse. In this project, syllabic-length acoustic models are being explored for the recognition of conversational speech. The project entails: definition of the acoustic features derived from energy trajectories spanning ca. 250-ms intervals of speech; statistical modeling of these syllable-length regions; development of a decoding scheme designed to combine the outputs of these acoustic models with the more traditional phone- and sub-phone-length models; and embedding the syllable, phone and sub-phone features into a multi-tiered representation of language designed for robust recognition under a wide range of acoustic-environmental and speaking conditions. A complete ASR system is being developed that will incorporate the results of this research, and will be evaluated on fluent speech. Successful results with the recognition system have the potential to improve practical ASR systems that must deal with the decoding of spontaneous discourse. Additionally, the analysis of longer-time structure at the acoustic, statistical, and lexical levels should improve our basic knowledge about the structure of conversation speech.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Collaborative Research: Towards Modeling Source Separation from Measured Cortical Responses
EAGER: Collaborative Research: Towards Modeling Human Speech Confusions in Noise
International: An Analysis of Speaker Diarization Systems Errors
CI-P: Towards a Consensus Representation for Understanding Structure of Multiparty Conversations
海外基金