III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
批准号:
1147810
负责人:
Eduard Hovy
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2013-02-28
中文摘要
帮助用户找到相关信息无疑是一个重要的问题,对当今以信息为基础的社会的运作至关重要。因此,全世界每天都有数百万人使用搜索引擎技术也就不足为奇了。尽管现有的搜索技术运行良好,但仍有相当大的改进空间。搜索引擎的创新是由快速、重复地测量给定系统产生的结果质量的能力所驱动的。这种类型的测量通常需要某种形式的人工输入。例如,可以聘请人类专家来评估搜索结果的相关性,或者搜索引擎可以记录用户交互,例如输入的查询和单击的结果。在收集了足够大的数据量之后,它就可以用来准确地衡量搜索引擎的质量。它还可以通过一个被称为“调整”或“训练”的过程来提高现有搜索引擎的质量。然而,收集大量此类信息通常需要大量的人力或计算资源。因此,持续的创新只有在非常高的成本下才能实现。构建无需人工操作的大型信息检索测试集的技术是本研究的主要焦点。而不是依赖于人工策划的信息,从Web中挖掘隐含的相关信号来自动构建大型的、可重用的测试集合,用于各种搜索任务,包括Web搜索、新闻搜索和企业搜索。观察到网络中包含大量的隐式关联信号是本研究的出发点。隐式相关性信号的最简单示例是超链接,源作者可以将其解释为确认目标页面相关性的信号。本研究探讨了这种隐式相关信号可以以完全无监督的方式有效挖掘和聚合以创建测试集合的假设,而无需任何人工努力。自动生成的测试集合以两种不同的方式进行评估。首先,测试集合是根据它们与人类生成的测试集合相比准确度量搜索系统质量的能力来评估的。其次,将使用自动测试集合调优的搜索引擎的质量与使用手动测试集合调优的引擎进行比较。这个项目更广泛的影响来自于自动构建的测试集合,这些测试集合自由地分发给更广泛的研究社区。由于在工业和学术环境中系统地评估和调整搜索引擎的训练数据的可用性增加,预计搜索引擎技术将取得进展。预计研究生和本科阶段的研究和教育的整合,以及通过各种外展项目吸引女性和代表性不足的学生,将产生更广泛的影响。
英文摘要
Helping users find relevant information is undeniably an important problem vital to the functioning of today's information-based societies. It is therefore no surprise that millions of people worldwide make use of search engine technologies each and every day. Although existing search technologies work well, there is still considerable room for improvement. Search engine innovation is driven by the ability to rapidly, and repeatedly, measure the quality of the results produced by a given system. This type of measurement typically requires some form of human input. For example, a human expert may be hired to assess the relevance of search results, or the search engine may log user interactions, such as the queries entered and the results clicked. After a sufficiently large amount of data has been collected, it can then be used to accurately measure search engine quality. It can also be used to improve the quality of existing search engines via a process known as "tuning" or "training". However, gathering large amounts of this information typically requires a significant amount of human effort or computational resources. Therefore, sustained innovation is only possible at a very steep cost.Techniques for constructing large information retrieval test collections that require no human effort are the primary focus of this research study. Rather than relying on human-curated information, implicit relevance signals from the Web are mined to automatically construct large, reusable test collections for a variety of search tasks, including Web search, news search, and enterprise search. The observation that the Web contains a large number of implicit relevance signals is the starting point of the research. The simplest example of an implicit relevance signal is the hyperlink, which can be interpreted as a signal acknowledging the relevance of the target page by the source author. The hypothesis that such implicit relevance signals can be effectively mined and aggregated in a completely unsupervised manner to create test collections without any human effort is investigated in this research. Automatically generated test collections are evaluated in two different ways. First, the test collections are evaluated according to their ability to accurately measure the quality of search systems compared to human-generated test collections. Second, the quality of search engines tuned using the automated test collections are compared against engines tuned using manual test collections.The broader impact of this project is derived from automatically constructed test collections that are freely distributed to the broader research community. Advances in search engine technologies are expected as the result of increased availability of training data to systematically evaluate and tune search engines, both in industrial and academic settings. Additional broader impact is expected from the integration of research and education at both the graduate and undergraduate levels and from engaging women and underrepresented students through various outreach programs.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: A Method to Retrieve Non-Textual Data from Widespread Repositories
-
批准号:1450545
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2014
-
负责人:Eduard Hovy
-
依托单位:
III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
-
批准号:1304939
-
项目类别:Standard Grant
-
资助金额:$10.32万
-
财政年份:2012
-
负责人:Eduard Hovy
-
依托单位:
EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
-
批准号:1265301
-
项目类别:Standard Grant
-
资助金额:$23.61万
-
财政年份:2012
-
负责人:Eduard Hovy
-
依托单位:
EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
-
批准号:1143703
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2011
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research III-COR: From a Pile of Documents to a Collection of Information: A Framework for Multi-Dimensional Text Analysis
-
批准号:0705091
-
项目类别:Standard Grant
-
资助金额:$32.0万
-
财政年份:2007
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research: Language Processing Technology for Electronic Rulemaking
-
批准号:0429360
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Eduard Hovy
-
依托单位:
Automating the Integration of EPA Databases
-
批准号:0306899
-
项目类别:Continuing Grant
-
资助金额:$90.0万
-
财政年份:2003
-
负责人:Eduard Hovy
-
依托单位:
SGER COLLABORATIVE: A Testbed for eRulemaking Data
-
批准号:0328175
-
项目类别:Standard Grant
-
资助金额:$2.5万
-
财政年份:2003
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research:Interlingual Annotation of Multilingual Text Corporation
-
批准号:0325021
-
项目类别:Standard Grant
-
资助金额:$16.88万
-
财政年份:2003
-
负责人:Eduard Hovy
-
依托单位:
ITR: Information Discovery in Digital Government: Self-extending Topic Maps and Ontologies (GrowOnto)
-
批准号:0205111
-
项目类别:Continuing Grant
-
资助金额:$100.0万
-
财政年份:2002
-
负责人:Eduard Hovy
-
依托单位:
Digital Government: dg.o Workshop and Publicity
-
批准号:0089522
-
项目类别:Continuing Grant
-
资助金额:$74.92万
-
财政年份:2000
-
负责人:Eduard Hovy
-
依托单位:
Workshop: Support for Workshop on Multilingual Information Management
-
批准号:9807199
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:1998
-
负责人:Eduard Hovy
-
依托单位:
Workshop: MT Summit Conference Support
-
批准号:9725058
-
项目类别:Standard Grant
-
资助金额:$0.5万
-
财政年份:1997
-
负责人:Eduard Hovy
-
依托单位:
International Language Generation Workshop, June 1994, Kennebunkport, ME
-
批准号:9321870
-
项目类别:Standard Grant
-
资助金额:$1.18万
-
财政年份:1994
-
负责人:Eduard Hovy
-
依托单位:
海外基金