课题基金 / 基金详情

RI: Medium: Collaborative Research: Learning Representations of Language for Domain Adaptation

RI: Medium: Collaborative Research: Learning Representations of Language for Domain Adaptation
RI:媒介:协作研究:学习领域适应的语言表示
批准号:
1065397
负责人:
Alexander Yates
金额:
$69.8万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-04-01 至 2016-03-31

项目摘要

项目成果

Alexander Yates的其他基金

相似基金

相关文献

中文摘要
翻译
监督自然语言处理(NLP)系统在与训练文本不同的领域和词汇上表现不佳。越来越多的经验和理论工作指出,传统自然语言处理系统使用的特征是领域依赖的罪魁祸首,并且无法概括为以前未见过的单词。这个项目是第一个系统地研究表征学习作为一种提高领域适应性能的技术。它探索了潜变量语言模型?包括阶乘隐马尔可夫模型、依赖关系解析模型和深层体系结构?作为从文本中提取新特征的技术。所得到的表示为分布相似的单词产生相似的特征,从而允许对在分类器的训练期间看不到的单词进行泛化。该项目还探索了培训语言模型的新程序,其中包括网络规模的ngram统计作为非监督培训中使用的标准统计的替代。语言用户具有非凡的创造力,新的话语领域不断出现,例如在专门的科学和技术领域。通过建立在这个项目产生的表示的基础上,NLP系统可以在新的领域和Web文本上提高准确性,使像语义网这样的应用更接近现实。对于资源匮乏的语言和领域,该项目可以通过减少对培训文本的广泛覆盖范围的需要,帮助降低文本注释的成本。通过让坦普尔大学和费城地区高中的不同学生团体参与,该项目有助于扩大代表不足的群体对计算机科学研究的参与。
英文摘要
Supervised Natural Language Processing (NLP) systems perform poorly on domains and vocabulary that differ from training texts. A growing body of empirical and theoretical work points to the features used by traditional NLP systems as the culprit for domain-dependence and for the inability to generalize to previously unseen words. This project is the first to systematically investigate representation-learning as a technique for improving performance on domain adaptation. It explores latent-variable language models ? including Factorial Hidden Markov Models, dependency parsing models, and deep architectures ? as techniques for extracting novel features from text. The resulting representations yield similar features for distributionally-similar words, thereby allowing generalization to words not seen during training of a classifier. The project also explores novel procedures for training a language model, which incorporate Web-scale ngram statistics as substitutes for standard statistics used in unsupervised training.Language users are extraordinarily inventive, and new domains of discourse appear constantly, such as in specialized areas of science and technology. By building on top of the representations produced by this project, NLP systems can improve in accuracy on new domains and on Web text, bringing applications like the Semantic Web closer to reality. For resource-poor languages and domains, the project can help reduce the cost of annotating texts by reducing the need for broad coverage in the training texts. By involving the diverse student bodies at Temple University and Philadelphia-area high schools, the project helps to broaden participation in computer science research by underrepresented groups.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Learning Open Domain Semantic Parsers
  • 批准号:
    1218692
  • 项目类别:
    Standard Grant
  • 资助金额:
    $42.58万
  • 财政年份:
    2012
  • 负责人:
    Alexander Yates
  • 依托单位:
海外基金