RI: Medium: Collaborative Research: Explicit Articulatory Models of Spoken Language, with Application to Automatic Speech Recognition
RI: Medium: Collaborative Research: Explicit Articulatory Models of Spoken Language, with Application to Automatic Speech Recognition
批准号:
0905633
负责人:
Karen Livescu
金额:
$43.88万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-07-01 至 2013-06-30
中文摘要
该奖项是根据2009年美国复苏和再投资法案(公法111-5)资助的。自动语音识别的主要挑战之一是说话风格的变化,包括语速变化和协同发音。发音器的模型(如嘴唇和舌头)可以简洁地表示这种可变性。以前关于发音模型的大部分工作都集中在声学和发音之间的关系上,但更重要的改进需要隐藏的发音状态结构的模型。这项工作既有提高识别率的技术目标,也有更好地理解发音现象的科学目标。该项目考虑了比以前研究的更大的模型类。特别是,该项目开发了图形模型,包括动态贝叶斯网络和条件随机场,旨在利用发音知识。为了认识到定向和非定向模型的益处,以及生成性和歧视性训练的益处,正在开发一种混合定向和非定向图形模型的新框架。该项目的活动包括通过语境建模、异步结构和专门训练对早期发音模型进行主要扩展;开发发音变量的因子化条件随机场模型;以及减轻单词混淆的区分训练。科学目标解决了发音轨迹在不同语境中如何变化的问题。使用了现有的数据库,并正在扩展人工发音注释的初步工作。此外,该项目使用发音模型来执行更大数据集的强制转录,为研究社区提供了额外的资源。其他广泛的影响包括适用于其他时间序列建模问题的新模型和新技术。扩大语音识别的适用性将有助于它实现其承诺,即实现更有效的语音信息存储和访问,并为听力或运动残疾人士提供平等的技术竞争环境。
英文摘要
This award is funded under the American Recovery and Reinvestment Act of 2009 (Public Law 111-5).One of the main challenges in automatic speech recognition is variability in speaking style, including speaking rate changes and coarticulation. Models of the articulators (such as the lips and tongue) can succinctly represent much of this variability. Most previous work on articulatory models has focused on the relationship between acoustics and articulation, but more significant improvements require models of the hidden articulatory state structure. This work has both a technological goal of improving recognition and a scientific goal of better understanding articulatory phenomena.The project considers larger model classes than previously studied. In particular, the project develops graphical models, including dynamic Bayesian networks and conditional random fields, designed to take advantage of articulatory knowledge. A new framework for hybrid directed and undirected graphical models is being developed, in recognition of the benefits of both directed and undirected models, and of both generative and discriminative training. The project activities include major extension of earlier articulatory models with context modeling, asynchrony structures, and specialized training; development of factored conditional random field models of articulatory variables; and discriminative training to alleviate word confusability.The scientific goal addresses questions about the ways in which articulatory trajectories vary in different contexts. Existing databases are used, and initial work in manual articulatory annotation is being extended. In addition, the project uses articulatory models to perform forced transcription of larger data sets, providing an additional resource for the research community. Other broad impacts include new models and techniques with applicability to other time-series modeling problems. Extending the applicability of speech recognition will help it fulfill its promise of enabling more efficient storage of and access to spoken information, and equalizing the technological playing field for those with hearing or motor disabilities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: From acoustics to semantics: Embedding speech for a hierarchy of tasks
-
批准号:1816627
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2018
-
负责人:Karen Livescu
-
依托单位:
EAGER: Discovery of Segmental Sub-Word Structure in Speech
-
批准号:1433485
-
项目类别:Standard Grant
-
资助金额:$9.99万
-
财政年份:2014
-
负责人:Karen Livescu
-
依托单位:
RI: Medium: Collaborative Research: Models of Handshape Articulatory Phonology for Recognition and Analysis of American Sign Language
-
批准号:1409837
-
项目类别:Standard Grant
-
资助金额:$85.41万
-
财政年份:2014
-
负责人:Karen Livescu
-
依托单位:
RI: Small: Multi-View Learning of Acoustic Features for Speech Recognition Using Articulatory Measurements
-
批准号:1321015
-
项目类别:Continuing Grant
-
资助金额:$44.49万
-
财政年份:2013
-
负责人:Karen Livescu
-
依托单位:
海外基金