课题基金 / 基金详情

CAREER: Discriminative Syntactic Language Modeling: Automatic Feature Selection and Efficient Annotation

CAREER: Discriminative Syntactic Language Modeling: Automatic Feature Selection and Efficient Annotation
职业:判别式句法语言建模:自动特征选择和高效注释
批准号:
0447214
负责人:
Brian Roark
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-04-01 至 2011-03-31

项目摘要

项目成果

Brian Roark的其他基金

相似基金

相关文献

中文摘要
翻译
这一建议的重点是在用于自动语音识别的语言建模的区别方法中有效地使用分析器派生的特征和标记器派生的特征。区分语言建模方法在定义特征方面提供了极大的灵活性,但潜在的语法分析器派生的特征空间的大小需要有效的特征标注和选择算法。该项目有四个具体目标。第一个目标是开发一套高效、通用和可扩展的句法特征选择算法,用于各种类型的标注和几种参数估计技术。第二个目标是开发通用的树和语法转换算法,旨在保留选定的特征注释,同时导致更快的解析甚至标记近似解析。第三个目标是评估在大词汇量连续语音识别(LVCSR)任务(即Switchboard)上使用的各种特征选择和语法转换方法。最终目标是设计和打包算法,以直接支持未来对其他应用程序的研究,如机器翻译(MT);以及其他语言,如中文和阿拉伯语。作为该项目的一部分开发的算法预计将有助于提高LVCSR的准确性和依赖这项技术的应用。这些算法被打包到一个公开可用的软件库中,使从事包括LVCSR和MT在内的许多应用领域以及各种语言的研究人员能够研究针对其特定任务的句法语言建模的最佳实践,而不必手动选择和评估特征集。
英文摘要
The focus of this proposal is on the effective use of parser-derived and tagger-derived features within discriminative approaches to language modeling for automatic speech recognition. Discriminativelanguage modeling approaches provide a tremendous amount of flexibility in defining features, but the size of the potential parser-derived feature space requires efficient feature annotation and selection algorithms. The project has four specific aims. The first aim is to develop a set of efficient, general, and scalable syntactic feature selection algorithms for use with various kinds of annotation and several parameter estimation techniques. The second aim is to develop general tree and grammar transformation algorithms designed to preserve selected feature annotations yet lead to faster parsing or even tagging approximations to parsing. The third aim is to evaluate a broad range of feature selection and grammar transformation approaches on a large vocabulary continuous speech recognition (LVCSR) task, namely Switchboard. The final aim is to design and package the algorithms to straightforwardly support future research into other applications, such as machine translation (MT); and into other languages, such as Chinese and Arabic. The algorithms developed as a part of this project are expected to contribute to improvements in LVCSR accuracy and applications that rely upon this technology. The algorithms are being packaged into a publicly available software library, enabling researchers working in many application areas including LVCSR and MT and various languages to investigate best practices in syntactic language modeling for their specific task, without having to hand-select and evaluate feature sets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Student Research Workshop in Computational Linguistics at the ACL 2009 Conference
RI-Small: Efficient hidden structure annotation via structural multiple-sequence alignments
SGER: RI: Text-Based Discriminative Language Modeling
海外基金