课题基金 / 基金详情

EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud

EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
EAGER:在云中构建、索引和搜索超级丰富的文档表示
批准号:
1143703
负责人:
Eduard Hovy
金额:
$25.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2012-12-31

项目摘要

项目成果

Eduard Hovy的其他基金

相似基金

相关文献

中文摘要
翻译
世界上每天都有数十亿新的数字文档被创造出来。例子包括电子邮件、博客文章、法律文件和新闻文章。为了实现有效的信息管理,许多这样的文档由信息检索系统处理,例如桌面搜索工具或Web搜索引擎。大多数现有技术以数字方式表示文档。对于计算机来说,这些表示只不过是一串比特,完全没有任何明确的含义。由于大多数现代搜索引擎使用这样的基本表示,它们经常不能正确地解释文档中找到的单词的含义,从而降低了结果的质量。尽管这个基本问题很重要,但令人惊讶的是,很少有人尝试构建并随后搜索编码文本丰富含义的文档表示,特别是对于包含数百万或数十亿个文本文档的数据集。本研究探讨了如何自动构建、索引和搜索下一代超级丰富的文档表示。该方法依赖于传统文本表示与基于自然语言处理的源(例如,命名实体、同义词和释义)、丰富的知识源(例如,Wikipedia和Freebase)、上下文源和其他增值的内容源的仔细集成。为大型文档集合构造这样的表示需要进行计算密集型的批处理,以便跨不同的数据源挖掘、聚合和连接数据。为了克服这些挑战,采用了可伸缩的大规模分布式云计算解决方案。生成的丰富的文档表示可以有效地应用于各种信息检索、自然语言处理和数据挖掘任务。
英文摘要
There are billions of new digital documents created around the world every day. Examples include emails, blog posts, legal documents, and news articles. To enable effective information management, many of these documents are processed by information retrieval systems, such as desktop search tools or Web search engines. Most existing technologies represent documents digitally. To a computer, these representations are nothing more than a sequence of bits, completely devoid of any explicit meaning. Since most modern search engines utilize such basic representations, they often fail to properly account for the meaning of the words found in the documents, thereby diminishing the quality of their results. Despite the importance of this fundamental problem, there have been surprisingly few attempts to build, and subsequently search, document representations that encode the deeply rich meaning of text, especially for data sets that contain millions or billions of text documents.This research investigates how to automatically construct, index, and search next-generation super-enriched document representations. The approach relies on the careful integration of traditional text representations with natural language processing-based sources (e.g., named entities, synonyms, and paraphrases), rich knowledge sources (e.g., Wikipedia and Freebase), contextual sources, and other value-added sources of content. Constructing such representations for large document collections requires computationally intensive batch processing to mine, aggregate, and join data across disparate sources. To overcome these challenges, a scalable, massively distributed cloud computing solution is adopted. The resulting enriched document representations can be effectively applied to a wide variety of information retrieval, natural language processing, and data mining tasks.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: A Method to Retrieve Non-Textual Data from Widespread Repositories
  • 批准号:
    1450545
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2014
  • 负责人:
    Eduard Hovy
  • 依托单位:
III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
  • 批准号:
    1304939
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.32万
  • 财政年份:
    2012
  • 负责人:
    Eduard Hovy
  • 依托单位:
EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
  • 批准号:
    1265301
  • 项目类别:
    Standard Grant
  • 资助金额:
    $23.61万
  • 财政年份:
    2012
  • 负责人:
    Eduard Hovy
  • 依托单位:
III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
  • 批准号:
    1147810
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2011
  • 负责人:
    Eduard Hovy
  • 依托单位:
海外基金