Building an Entity-Centric Stream Filtering Test Collection for TREC 2012

Building an Entity-Centric Stream Filtering Test Collection for TREC 2012
复制标题

DOI:
--
复制
发表时间:
2012-11
期刊:
--
影响因子:
--
通讯作者:
John R. Frank;Max Kleiman-Weiner;D. Roberts;Feng Niu;Ce Zhang;Christopher Ré;I. Soboroff
John R. Frank;Max Kleiman-Weiner;D. Roberts;Feng Niu;Ce Zhang;Christopher Ré;I. Soboroff
中科院分区:
其他
文献类型:
--
作者:
John R. Frank;Max Kleiman-Weiner;D. Roberts;Feng Niu;Ce Zhang;Christopher Ré;I. Soboroff

文献摘要

被引文献

相似文献

摘要:TREC 2012中的知识库加速轨道专注于一个单一任务:过滤与预定义实体列表高度相关的按时间排序的语料库。KBA与以前的过滤评估在两个主要方面有所不同:流语料库比以前的过滤集合大100倍,并且使用实体作为主题使系统能够将结构化知识库(KB),如Wikipedia,作为外部数据源。一个成功的KBA系统必须做的不仅仅是通过将文档链接到知识库来解决实体提及的含义:它还必须区分在实体的WP文章中值得引用的集中相关文档。这结合了自然语言处理(NLP)和信息检索(IR)的思维。TREC中的过滤轨道通常使用基于由一组关键字查询或简短描述描述的主题的查询,注释者根据他们对主题的个人解释生成相关性判断。对于TREC 2012,我们选择了一组基于维基百科实体的过滤主题:27个人和2个组织。这种命名实体在NLP中比IR更熟悉。我们还构建了一个全新的流语料库,从2011年10月到2012年4月,跨越4,973个连续小时。它包含超过400M的文档,我们为标识为英语的40%的文档添加了命名实体分类标签。每个文档都有一个将其放置在流中的时间戳。29个目标实体在语料库中被提及的频率不够高,以至于NIST评估人员可以判断大多数提及文档的相关性(91%)。2012年1月之前的文件判断作为培训数据提供给TREC团队,用于过滤剩余时间的文件。根据评估员生成的值得引用的文档列表对运行提交进行评估。我们给出了所有运行提交的实体的平均F_1分数峰值。高分系统
Abstract : The Knowledge Base Acceleration track in TREC 2012 focused on a single task: filter a time-ordered corpus for documents that are highly relevant to a predefined list of entities. KBA differs from previous filtering evaluations in two primary ways: the stream corpus is 100x larger than previous filtering collections, and the use of entities as topics enables systems to incorporate structured knowledge bases (KB), such as Wikipedia, as external data sources. A successful KBA system must do more than resolve the meaning of entity mentions by linking documents to the KB: it must also distinguish centrally relevant documents that are worth citing in the entity's WP article. This combines thinking from natural language processing (NLP) and information retrieval (IR). Filtering tracks in TREC have typically used queries based on topics described by a set of keyword queries or short descriptions, and annotators have generated relevance judgments based on their personal interpretation of the topic. For TREC 2012, we selected a set of filter topics based on Wikipedia entities: 27 people and 2 organizations. Such named entities are more familiar in NLP than IR. We also constructed an entirely new stream corpus spanning 4,973 consecutive hours from October 2011 through April 2012. It contains over 400M documents, which we augmented with named entity classification tagging for the 40% of the documents identified as English. Each document has a timestamp that places it in the stream. The 29 target entities were mentioned infrequently enough in the corpus that NIST assessors could judge the relevance of most of the mentioning documents (91%). Judgments for documents from before January 2012 were provided to TREC teams as training data for filtering documents from the remaining hours. Run submissions were evaluated against the assessor-generated list of citation-worthy documents. We present peak F_1 scores averaged across the entities for all run submissions. High scoring system