STIMULATE: Exploiting Nonlocal and Syntactic Word Relationships in Language Models for Conversational Speech Recognition
STIMULATE: Exploiting Nonlocal and Syntactic Word Relationships in Language Models for Conversational Speech Recognition
批准号:
9618874
负责人:
Frederick Jelinek
金额:
$75.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-03-01 至 2001-02-28
中文摘要
能够以语音和手写等方式与人类用户交互的自动化系统将大大提高生产力和系统可用性。这些系统将使人们能够简单地访问互联网上的信息和服务。这些功能对于其他任务至关重要,例如允许残疾用户访问或在执行复杂的维修时查询在线维护手册。语言的统计模型是这类系统的关键组成部分,它在语音或手写与文本之间进行转换,在统计机器翻译系统中也是如此。目前大多数语言建模算法都表现出严重的短视,它们对下一个单词的预测只基于前面的几个单词。当人类面对一个类似的任务时,他们很容易胜过这些模型,因为他们可以从更完整的语境中获得更丰富的语言信息。CLSP的研究人员建议研究和开发新的语言建模技术,利用更丰富的上下文信息。他们建议检查使用各种技术的模型,通过主题的动态、分层模型来捕获语法依赖性,并使用最大熵原理将结果模型与当前最好的模型结合起来。本研究将专注于提高人类自发语音的识别准确性,但除此之外,还将提供适用于所有语言建模应用的新信息源和技术的见解。
英文摘要
Automated systems that can interact with human users in modalities such as speech and handwriting will greatly enhance productivity and system usability. These systems will allow simple access to information and services on the internet. These capabilities are essential to other tasks, such as enabling access by handicapped users or querying an on-line maintenance manual while performing intricate repairs. A statistical model of language is a crucial component in such systems, which convert between speech or handwriting and text, and in statistical machine translation systems. Most current algorithms for language modeling exhibit an acute myopia, basing their predictions of the next word on only a few immediately preceding words. When humans are faced with a comparable task they easily outperform these models using the richer linguistic information available to them from more complete context. Researchers at CLSP propose to investigate and develop novel language modeling techniques that exploit richer contextual information. They propose to examine models that use a variety of techniques to capture syntactic dependencies through dynamic, hierarchical models of topic, and to combine the resulting models with the best current ones using the maximum entropy principle. This research will focus on improving the recognition accuracy of spontaneous human speech, but beyond that will provide insight into new information sources and techniques applicable to all applications of language modeling.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Cross-Cutting Research Workshops in Intelligent Information Systems
-
批准号:0833652
-
项目类别:Standard Grant
-
资助金额:$28.05万
-
财政年份:2009
-
负责人:Frederick Jelinek
-
依托单位:
RI: Cross-Cutting Research Workshops in Intelligent Information Systems
-
批准号:0705708
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Frederick Jelinek
-
依托单位:
2001 Language Engineering Workshop for Students and Professionals: Integrating Research and Education (WS01)
-
批准号:0097467
-
项目类别:Standard Grant
-
资助金额:$58.3万
-
财政年份:2001
-
负责人:Frederick Jelinek
-
依托单位:
Robust Knowledge Discovery from Parallel Speech and Text Sources
-
批准号:9982329
-
项目类别:Continuing Grant
-
资助金额:$54.64万
-
财政年份:2001
-
负责人:Frederick Jelinek
-
依托单位:
ITR/IM+PE+SY: Summer Workshops on Human Language Technology: Integrating Research and Education
-
批准号:0121285
-
项目类别:Continuing Grant
-
资助金额:$234.95万
-
财政年份:2001
-
负责人:Frederick Jelinek
-
依托单位:
Workshop: 2000 Language Engineering Workshop for Students and Professionals: Integrating Research and Education (WS00)
-
批准号:0071215
-
项目类别:Standard Grant
-
资助金额:$75.02万
-
财政年份:2000
-
负责人:Frederick Jelinek
-
依托单位:
U.S.-Czech Cooperative Research: Speech Recognition of a Slavic Language
-
批准号:9810517
-
项目类别:Standard Grant
-
资助金额:$28.2万
-
财政年份:1999
-
负责人:Frederick Jelinek
-
依托单位:
1999 Language Engineering Workshop for Students and Professionals: Integrating Research and Education (WS99)
-
批准号:9820687
-
项目类别:Standard Grant
-
资助金额:$88.2万
-
财政年份:1999
-
负责人:Frederick Jelinek
-
依托单位:
WORKSHOP: Language Engineering Workshop for Students and Professionals Integrating Research and Education (WS98)
-
批准号:9732388
-
项目类别:Standard Grant
-
资助金额:$75.0万
-
财政年份:1998
-
负责人:Frederick Jelinek
-
依托单位:
海外基金