课题基金 / 基金详情

Collaborative Research: Landmark-based Robust Speech Recognition using Prosody-guided Models of Speech Variability

Collaborative Research: Landmark-based Robust Speech Recognition using Prosody-guided Models of Speech Variability
协作研究:使用韵律引导的语音变异模型进行基于地标的鲁棒语音识别
批准号:
0703805
负责人:
Abeer Alwan
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-06-01 至 2011-05-31

项目摘要

项目成果

Abeer Alwan的其他基金

相似基金

相关文献

中文摘要
翻译
尽管自动语音识别技术取得了长足的进步,但在自动转录不受限制的会话语音、代表许多说话者和方言、嵌入不利的声学环境方面,我们还没有一个性能可与人类媲美的系统。该方法采用新的高维机器学习技术,受语音产生和感知的经验和理论研究的限制,从数据中学习人类听众从语音中提取的信息结构。为了做到这一点,我们将通过推导反映人类语音产生和语音感知的知识表示来开发语音声学、发音可变性、韵律和语法的大词汇心理现实模型,使用机器学习技术同时调整所有知识表示的参数,以尽量减少识别器的结构风险。该团队将开发非线性声学地标检测器和模式分类器,集成基于听觉的信号处理和声学语音处理,对噪声、扬声器特性和混响的变化不影响,并且可以以半监督的方式从标记和未标记的数据中学习。此外,他们将使用可变帧率分析,这将允许多分辨率分析,以及实现基于手势的词法访问,使用各种训练数据。这项工作将改善人与机器之间的沟通和协作,也将提高对人类如何产生和感知语言的理解。这项工作汇集了语音处理,声学语音学,韵律学,手势音韵学,统计模式匹配,语言建模和语音感知方面的专家团队,以及工程,计算机科学和语言学的教师。支持和参与的学生和博士后研究员是项目的一部分,从事语音建模和算法开发。最后,拟议的工作将产生一套数据库和工具,这些数据库和工具将被分发给整个研究和教育界。
英文摘要
Proposal ID 0703859 Date 04/11/2007 Despite great strides in the development of automatic speech recognition technology, we do not yet have a system with performance comparable to humans in automatically transcribing unrestricted conversational speech, representing many speakers and dialects, and embedded in adverse acoustic environments. This approach applies new high-dimensional machine learning techniques, constrained by empirical and theoretical studies of speech production and perception, to learn from data the information structures that human listeners extract from speech. To do this, we will develop large-vocabulary psychologically realistic models of speech acoustics, pronunciation variability, prosody, and syntax by deriving knowledge representations that reflect those proposed for human speech production and speech perception, using machine learning techniques to adjust the parameters of all knowledge representations simultaneously in order to minimize the structural risk of the recognizer. The team will develop nonlinear acoustic landmark detectors and pattern classifiers that integrate auditory-based signal processing and acoustic phonetic processing, are invariant to noise, change in speaker characteristics and reverberation, and can be learned in a semi-supervised fashion from labeled and unlabeled data. In addition, they will use variable frame rate analysis, which will allow for multi-resolution analysis, as well as implement lexical access based on gesture, using a variety of training data. The work will improve communication and collaboration between people and machines and also improve understanding of how human produce and perceive speech. The work brings together a team of experts in speech processing, acoustic phonetics, prosody, gestural phonology, statistical pattern matching, language modeling, and speech perception, with faculty across engineering, computer science and linguistics. Support and engagement of students and postdoctoral fellows are part of the project, engaging in speech modeling and algorithm development. Finally, the proposed work will result in a set of databases and tools that will be disseminated to serve the research and education community at large.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Improving speech technology for better learning outcomes: the case of AAE child speakers
Collaborative Research: RI: Small: From Ultrasound and MRI to articulatory and acoustic models of child speech development
Workshop for Undergraduate and MS Female Students in Speech Science and Technology
NRI: INT: COLLAB: Development, Deployment and Evaluation of Personalized Learning Companion Robots for Early Literacy and Language Learning
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)