A Document Analysis System for Linking Cross-document Entities
A Document Analysis System for Linking Cross-document Entities
复制标题
DOI:
--
复制
发表时间:
2012-07
期刊:
影响因子:
--
通讯作者:
Manabu Ohta;A. Takasu
中科院分区:
文献类型:
--
作者:
Manabu Ohta;A. Takasu
—This paper proposes an entity extraction and matching system for digital documents. Digital documents usually contain many links to their relevant information, but they do not cover all the links. Entity extraction and matching systems are used to detect such implicit links. They usually consist of several steps such as parsing, dictionary matching, and classification. Some of these steps, however, inevitably cause errors, which must be managed properly so that the process of subsequent steps is not degraded. We have therefore been developing an entity extraction and matching system focusing on managing the errors incurred at each step. This paper overviews the system and explains some techniques we have developed to improve the quality of entity extraction and matching because the system can be a key solution to content management for institutional repositories and academic societies as well as digital libraries.