课题基金 / 基金详情

CAREER: Resolving Lexical Ambiguities in Natural Language Processing

CAREER: Resolving Lexical Ambiguities in Natural Language Processing
职业:解决自然语言处理中的词汇歧义
批准号:
9985033
负责人:
David Yarowsky
金额:
$33.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2000
资助国家:
美国
项目状态:
已结题
起止时间:
2000-07-01 至 2006-06-30

项目摘要

项目成果

David Yarowsky的其他基金

相似基金

相关文献

中文摘要
翻译
这是为期4年的持续资助奖的第一年。解决自然语言中固有的歧义问题是人与机器高效准确交流的主要障碍之一。开发解决方案的一个主要瓶颈是严重缺乏区分词义的训练数据,以及手动输入这些信息的高昂成本。该项目的一个重点是开发无监督和最少监督的算法,以便在不需要昂贵的手工标记训练数据的情况下获得这种技能。这些方法将利用在非常大的文本语料库(超过100亿字)中观察到的分布特性;PI还将研究更丰富的特征空间表示、类模型、平滑方法和专门用于在超高维特征空间中分类的学习算法。除了词义消除歧义的问题外,该项目还将探索一系列密切相关的词汇歧义任务的解决方案,包括拼写更正、专有名称分类、大写恢复、重音和发音恢复、多种语言、希伯来语和阿拉伯语的元音恢复、同形异义词的语音合成、机器翻译中的词汇选择以及语音识别中在语音易混淆的词典中选择的某些方面。这些不同的问题通常不被认为是同一个班级的成员,这个项目试图通过开发关于班级中一个成员的方法和培训数据,并利用关于班级中其他问题的方法和数据来利用存在的协同效应。因此,这种统一的方法为人机交互和信息提取中的关键问题提供了快速并行进展的潜力。
英文摘要
This is the first year of funding of a 4 year, continuing award. One of the major roadblocks in theefficient and accurate communication between humans and machines is the resolution of theambiguity inherent in natural languages. A major bottleneck in developing solutions is the severeshortage of training data that distinguishes word senses, and the high cost of inputting thisinformation manually. A focus of this project is the development of unsupervised and minimallysupervised algorithms for acquiring such skills without costly hand-tagged training data. Suchmethods will exploit the distribution properties observed in very large text corpora (over 10 billionwords); the PI will also investigate richer representations of feature space, class models,smoothing methods and learning algorithms specialized for classification in very high-dimensionalfeaturespaces. In addition to the problem of word-sense disambiguation, this project will exploreshared solutions to a closely related set of lexical-ambiguity tasks including spelling correction,propername classification, capitalization restoration, accent and diacritic restoration for, diverselanguages, vowel restoration in Hebrew and Arabic, speech synthesis on homographs, lexicalchoice in machine translation, and certain aspects of choosing among phonetically confusablecandidates in speech recognition. These diverse problems are not normally recognized as beingmembers of the same class, and this project seeks to exploit the synergies present by developingmethods and training data on one member of the class and utilizing the methods and data on other.problems in the class. Thus this unified approach offers the potential for rapid parallel progress onkey problems in human-computer interaction and information extraction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Scientific Community On-Site Assessment Workshop for Robust Intelligence
  • 批准号:
    0839056
  • 项目类别:
    Standard Grant
  • 资助金额:
    $5.0万
  • 财政年份:
    2008
  • 负责人:
    David Yarowsky
  • 依托单位:
RI: Multi-Level Modeling of Language and Translation
  • 批准号:
    0713448
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.12万
  • 财政年份:
    2007
  • 负责人:
    David Yarowsky
  • 依托单位:
海外基金