课题基金 / 基金详情

ITR: Framenet++: An On-Line Lexical Semantic Resource and its Application to Speech and Language Understanding

ITR: Framenet++: An On-Line Lexical Semantic Resource and its Application to Speech and Language Understanding
ITR:框架网:在线词汇语义资源及其在语音和语言理解中的应用
批准号:
0086132
负责人:
Charles Fillmore
金额:
$209.84万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2000
资助国家:
美国
项目状态:
已结题
起止时间:
2000-09-01 至 2004-08-31

项目摘要

项目成果

Charles Fillmore的其他基金

相似基金

相关文献

中文摘要
翻译
这是为期三年的持续奖励的第一年资助。健壮的独立于领域的语言理解对于多语言信息提取、摘要、问答和自动翻译是必不可少的。随着普适计算环境即将到来,语言理解将变得更加不可或缺,以便与功能迥异的构件进行交互。在过去的15年里,自然语言理解领域取得了重大进展。这一收益在很大程度上归功于统计算法与基于模板的算法的复杂组合,这些算法是为空中交通信息、旅行计划和商业新闻等特定领域量身定做的。但要真正解决与领域无关的理解问题,就需要从基于模板的单语系统转向更灵活、更通用的人机界面系统,这需要通过三个关键创新:(1)作为这些理解系统的后端的与领域无关的语义语言,取代现有的受限于领域的模板和槽;(2)丰富的语义词汇库,它足够广泛,足以涵盖语言工程任务所需的单词,并且足够深入的可用语义信息,以支持真正的领域无关理解;这个项目将开发这三个组件:一个非常大的词汇库FrameNet++,一种为独立于领域的理解任务而设计的语义语言,以及将其应用于关键NLU应用程序并对其进行评估的工具。语义语言和词汇数据库的基础是形式化的语义框架以及英语词典中很大一部分的语义和句法组合属性--配价。FrameNet++将通过描述定义单词的概念框架并确定这些单词的论元可以承担的语义角色,提供比当前数据库(如Comlex和WordNet)中可用的更丰富的语义信息。这些角色和框架是构建独立于领域的语言理解应用程序的关键。该项目将从一开始就将重点放在具体的自然语言理解应用上:词义消除歧义、信息提取、多语言信息提取,以及最终扩展到文本数据挖掘。对于每个应用程序,PI和他的团队将应用FrameNet++系统来提高语义组件的领域独立性,使用我们已经开始实现的语义注释的统计算法。这些应用程序将反过来提供丰富和现实的评估框架来指导FrameNet++的开发,并将鼓励潜在用户将其应用于各种任务。FrameNet++数据库将能够服务于多种目的。它提供了关于词频、词/义映射以及与词义相关的组合模式的统计信息,将用于各种自动语言理解过程,包括词义消除歧义和信息提取。由于形式语义标注是独立于任何单独语言的概念结构的关键,因此它们可用于创建其他语言的平行词典数据库。数据库中的语义结构将促进机器翻译和机器辅助翻译中从一种语言到另一种语言的匹配,而句法结构则允许用目标语言产生适当的语法句子。
英文摘要
This is the first year funding of a three-year continuing award. Robust domain-independent language understanding is essential for multilingual information extraction, summarization, question answering, and automatic translation. With pervasive computing environments soon to come, language understanding will become even more indispensable for interacting with artifacts of widely different functionalities. The field of natural language understanding has made significant progress in the last fifteen years. A large part of this gain is due to the sophisticated combination of statistical algorithms with template-based algorithms tailored to specific domains like air-traffic information, travel scheduling, and business news. But any real solution to the problem of domain-independent understanding will require moving beyond template-based monolingual systems to more flexible, general purpose HCI systems via three key innovations: (1) a domain-independent semantic language as the back end for these understanding systems, replacing the current domain-restricted templates and slots; (2) rich semantic lexical databases which are broad enough to cover the necessary words for language engineering tasks, and deep enough in usable semantic information to support true domain-independent understanding; and (3) sophisticated techniques for performing this mapping.This project will develop these three components: a very large lexical database FrameNet++, a semantic language designed for domain-independent understanding tasks, and the tools for applying it to and evaluating it on key NLU applications. The semantic language and lexical database are based on formalizing the semantic frames and the semantic and syntactic combinatory properties - the valences - of a significant portion of the English lexicon. FrameNet++ will offer significantly richer semantic information than is available in current databases like COMLEX and WordNet, by characterizing the conceptual frames within which words are defined and identifying the semantic roles which the arguments of these words can take. These roles and frames are key to building domain-independent language understanding applications. The project will focus from the start on specific NLU applications: word sense disambiguation, information extraction, multilingual information extraction, and an eventual extension to text data mining. For each application, the PI and his team will apply the FrameNet++ system to improve the domain independence of the semantic components, using statistical algorithms for semantic annotation that we have already begun to implement. These applications will in turn provide a rich and realistic evaluation framework to guide FrameNet++ development, and will encourage potential users to apply it to a wide variety of tasks.The FrameNet++ database will be capable of serving many purposes. Provided with statistical information about frequencies of words, word/sense mappings, and combinatorial patterns linked to word senses, it will be usable in various automatic language understanding processes, including word sense disambiguation and information extraction. Since the formal semantic annotations are keyed to conceptual structures which are independent of any individual language, they are available for the creation of parallel lexicon databases of other languages. The semantic structures in the databases will facilitate matches from one language to another, in machine translation and machine-assisted translation, while the syntactic structures allow the production of appropriate grammatical sentences in the target language.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SGER: Beyond the Core: A Pilot Project on Cataloguing Grammatical Constructions and Multiword Expressions in English.
STIMULATE: Tools for Lexicon Building
Lexical Semantics and Deixis
  • 批准号:
    7503538
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.9万
  • 财政年份:
    1974
  • 负责人:
    Charles Fillmore
  • 依托单位:
海外基金