Terms over LOAD: Leveraging Named Entities for Cross-Document Extraction and Summarization of Events

Terms over LOAD: Leveraging Named Entities for Cross-Document Extraction and Summarization of Events
复制标题

DOI:
10.1145/2911451.2911529
复制
发表时间:
2016-07
期刊:
Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Andreas Spitz;Michael Gertz
Andreas Spitz;Michael Gertz
中科院分区:
其他
文献类型:
--
作者:
Andreas Spitz;Michael Gertz

文献摘要

被引文献

相似文献

真实的世界事件,如历史事件,通常包含空间和时间方面,并涉及一个特定的人群。这反映在文本来源中对事件的描述中,其中包含对命名实体和日期的提及。然而,给定大量的文档集合,这样的描述在单个文档中可能是不完整的,或者分散在多个文档中。在这些情况下,利用关于事件中涉及的实体的部分信息来提取缺失的信息是有益的。在本文中,我们介绍了负载模型的跨文档事件提取在大规模的文档集合。基于图的模型依赖于属于类位置、组织、参与者和日期的命名实体的共现,并将它们置于周围术语的上下文中。因此,该模型允许高效的查询,并且可以在可忽略的时间内增量地更新,以反映对底层文档集合的更改。我们讨论了这种方法的多功能性事件总结,完成部分事件信息,并提取命名实体和日期的描述。我们创建并提供一个加载图的文件在英语维基百科从命名的实体提取的国家的最先进的NER工具。基于包括不同事件摘要的历史数据的评估集,我们评估结果图。我们发现,该模型不仅允许近实时检索的信息从底层的文档集合,但也提供了一个全面的框架浏览和总结事件数据。
Real world events, such as historic incidents, typically contain both spatial and temporal aspects and involve a specific group of persons. This is reflected in the descriptions of events in textual sources, which contain mentions of named entities and dates. Given a large collection of documents, however, such descriptions may be incomplete in a single document, or spread across multiple documents. In these cases, it is beneficial to leverage partial information about the entities that are involved in an event to extract missing information. In this paper, we introduce the LOAD model for cross-document event extraction in large-scale document collections. The graph-based model relies on co-occurrences of named entities belonging to the classes locations, organizations, actors, and dates and puts them in the context of surrounding terms. As such, the model allows for efficient queries and can be updated incrementally in negligible time to reflect changes to the underlying document collection. We discuss the versatility of this approach for event summarization, the completion of partial event information, and the extraction of descriptions for named entities and dates. We create and provide a LOAD graph for the documents in the English Wikipedia from named entities extracted by state-of-the-art NER tools. Based on an evaluation set of historic data that include summaries of diverse events, we evaluate the resulting graph. We find that the model not only allows for near real-time retrieval of information from the underlying document collection, but also provides a comprehensive framework for browsing and summarizing event data.