EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
批准号:
1265301
负责人:
Eduard Hovy
金额:
$23.61万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-09-01 至 2014-08-31
中文摘要
世界各地每天都有数十亿份新的数字文档产生。示例包括电子邮件、博客文章、法律的文档和新闻文章。为了实现有效的信息管理,这些文件中有许多是由信息检索系统处理的,如桌面搜索工具或Web搜索引擎。大多数现有技术都以数字方式表示文档。对于计算机来说,这些表示只不过是一个比特序列,完全没有任何明确的意义。由于大多数现代搜索引擎都使用这种基本表示法,它们往往无法正确解释文档中单词的含义,从而降低了搜索结果的质量。尽管这个基本问题的重要性,有令人惊讶的是,很少有人尝试建立,并随后搜索,文档表示,编码的文本的深刻丰富的含义,特别是对于包含数百万或数十亿的文本documents.This研究的数据集,探讨如何自动构建,索引和搜索下一代超丰富的文档表示。该方法依赖于传统文本表示与基于自然语言处理的源(例如,命名实体、同义词和释义),丰富的知识源(例如,Wikipedia和Freebase)、上下文来源和其他增值内容来源。为大型文档集合构建这样的表示需要计算密集型批处理来挖掘、聚合和连接不同来源的数据。为了克服这些挑战,采用了可扩展的大规模分布式云计算解决方案。由此产生的丰富的文档表示可以有效地应用于各种各样的信息检索,自然语言处理和数据挖掘任务。
英文摘要
There are billions of new digital documents created around the world every day. Examples include emails, blog posts, legal documents, and news articles. To enable effective information management, many of these documents are processed by information retrieval systems, such as desktop search tools or Web search engines. Most existing technologies represent documents digitally. To a computer, these representations are nothing more than a sequence of bits, completely devoid of any explicit meaning. Since most modern search engines utilize such basic representations, they often fail to properly account for the meaning of the words found in the documents, thereby diminishing the quality of their results. Despite the importance of this fundamental problem, there have been surprisingly few attempts to build, and subsequently search, document representations that encode the deeply rich meaning of text, especially for data sets that contain millions or billions of text documents.This research investigates how to automatically construct, index, and search next-generation super-enriched document representations. The approach relies on the careful integration of traditional text representations with natural language processing-based sources (e.g., named entities, synonyms, and paraphrases), rich knowledge sources (e.g., Wikipedia and Freebase), contextual sources, and other value-added sources of content. Constructing such representations for large document collections requires computationally intensive batch processing to mine, aggregate, and join data across disparate sources. To overcome these challenges, a scalable, massively distributed cloud computing solution is adopted. The resulting enriched document representations can be effectively applied to a wide variety of information retrieval, natural language processing, and data mining tasks.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: A Method to Retrieve Non-Textual Data from Widespread Repositories
-
批准号:1450545
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2014
-
负责人:Eduard Hovy
-
依托单位:
III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
-
批准号:1304939
-
项目类别:Standard Grant
-
资助金额:$10.32万
-
财政年份:2012
-
负责人:Eduard Hovy
-
依托单位:
III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
-
批准号:1147810
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2011
-
负责人:Eduard Hovy
-
依托单位:
EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
-
批准号:1143703
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2011
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research III-COR: From a Pile of Documents to a Collection of Information: A Framework for Multi-Dimensional Text Analysis
-
批准号:0705091
-
项目类别:Standard Grant
-
资助金额:$32.0万
-
财政年份:2007
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research: Language Processing Technology for Electronic Rulemaking
-
批准号:0429360
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Eduard Hovy
-
依托单位:
Automating the Integration of EPA Databases
-
批准号:0306899
-
项目类别:Continuing Grant
-
资助金额:$90.0万
-
财政年份:2003
-
负责人:Eduard Hovy
-
依托单位:
SGER COLLABORATIVE: A Testbed for eRulemaking Data
-
批准号:0328175
-
项目类别:Standard Grant
-
资助金额:$2.5万
-
财政年份:2003
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research:Interlingual Annotation of Multilingual Text Corporation
-
批准号:0325021
-
项目类别:Standard Grant
-
资助金额:$16.88万
-
财政年份:2003
-
负责人:Eduard Hovy
-
依托单位:
ITR: Information Discovery in Digital Government: Self-extending Topic Maps and Ontologies (GrowOnto)
-
批准号:0205111
-
项目类别:Continuing Grant
-
资助金额:$100.0万
-
财政年份:2002
-
负责人:Eduard Hovy
-
依托单位:
Digital Government: dg.o Workshop and Publicity
-
批准号:0089522
-
项目类别:Continuing Grant
-
资助金额:$74.92万
-
财政年份:2000
-
负责人:Eduard Hovy
-
依托单位:
Workshop: Support for Workshop on Multilingual Information Management
-
批准号:9807199
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:1998
-
负责人:Eduard Hovy
-
依托单位:
Workshop: MT Summit Conference Support
-
批准号:9725058
-
项目类别:Standard Grant
-
资助金额:$0.5万
-
财政年份:1997
-
负责人:Eduard Hovy
-
依托单位:
International Language Generation Workshop, June 1994, Kennebunkport, ME
-
批准号:9321870
-
项目类别:Standard Grant
-
资助金额:$1.18万
-
财政年份:1994
-
负责人:Eduard Hovy
-
依托单位:
海外基金