Document Filtering for Long-tail Entities

Document Filtering for Long-tail Entities
复制标题

DOI:
10.1145/2983323.2983728
复制
发表时间:
2016-09
期刊:
Proceedings of the 25th ACM International on Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
R. Reinanda;E. Meij;M. de Rijke
R. Reinanda;E. Meij;M. de Rijke
中科院分区:
其他
文献类型:
--
作者:
R. Reinanda;E. Meij;M. de Rijke

文献摘要

被引文献

相似文献

在知识库建设和维护的背景下,过滤与实体相关的文档是一项必不可少的任务。它需要处理可能与一个实体有关的按时间排序的文件流,以便只选择那些包含重要信息的文件。用于流行实体的文档过滤的最新方法是实体依赖的:它们依赖于每个特定实体的区分特征的细节,并且还根据这些细节进行训练。此外,这些方法倾向于使用所谓的外部信息,例如维基百科页面浏览量和相关实体,这些信息通常仅适用于流行的头部实体。因此,基于这种信号的依赖于可靠性的方法不适合作为长尾实体的过滤方法。在本文中,我们提出了一个文件过滤方法长尾实体,是实体独立的,因此也推广到看不见或很少看到的实体。它基于内在特征,即,从提到实体的文档中导出的特征。我们提出了一套功能,捕捉信息量,实体显着性和及时性。特别是,我们引入功能的基础上,实体方面的相似性,关系模式,和时间的表达和联合收割机这些标准功能的文件过滤。根据TREC KBA 2014在公开数据集上的设置进行的实验表明,我们的模型能够在多个基线上提高长尾实体的过滤性能。将该模型应用于看不见的实体的结果是有希望的,这表明该模型能够学习重要文档的一般特征。所有实体的总体业绩-即,不仅仅是长尾实体-在不依赖于任何实体特定训练数据的情况下改进了现有技术。
Filtering relevant documents with respect to entities is an essential task in the context of knowledge base construction and maintenance. It entails processing a time-ordered stream of documents that might be relevant to an entity in order to select only those that contain vital information. State-of-the-art approaches to document filtering for popular entities are entity-dependent: they rely on and are also trained on the specifics of differentiating features for each specific entity. Moreover, these approaches tend to use so-called extrinsic information such as Wikipedia page views and related entities which is typically only available only for popular head entities. Entity-dependent approaches based on such signals are therefore ill-suited as filtering methods for long-tail entities. In this paper we propose a document filtering method for long-tail entities that is entity-independent and thus also generalizes to unseen or rarely seen entities. It is based on intrinsic features, i.e., features that are derived from the documents in which the entities are mentioned. We propose a set of features that capture informativeness, entity-saliency, and timeliness. In particular, we introduce features based on entity aspect similarities, relation patterns, and temporal expressions and combine these with standard features for document filtering. Experiments following the TREC KBA 2014 setup on a publicly available dataset show that our model is able to improve the filtering performance for long-tail entities over several baselines. Results of applying the model to unseen entities are promising, indicating that the model is able to learn the general characteristics of a vital document. The overall performance across all entities---i.e., not just long-tail entities---improves upon the state-of-the-art without depending on any entity-specific training data.