A Document Analysis System for Linking Cross-document Entities

A Document Analysis System for Linking Cross-document Entities
复制标题

DOI:
--
复制
发表时间:
2012-07
期刊:
--
影响因子:
--
通讯作者:
Manabu Ohta;A. Takasu
Manabu Ohta;A. Takasu
中科院分区:
其他
文献类型:
--
作者:
Manabu Ohta;A. Takasu

文献摘要

相似文献

本文提出了一个数字文档的实体提取和匹配系统。数字文档通常包含许多指向其相关信息的链接,但它们并不涵盖所有链接。实体提取和匹配系统被用来检测这种隐含的链接。它们通常包括几个步骤,如解析,字典匹配和分类。然而,这些步骤中的一些不可避免地会导致错误,这些错误必须得到适当的管理,以便后续步骤的过程不会降级。因此,我们一直在开发一个实体提取和匹配系统,专注于管理在每个步骤中发生的错误。本文概述了该系统,并解释了一些技术,我们已经开发,以提高实体提取和匹配的质量,因为该系统可以是一个关键的解决方案,机构知识库和学术团体以及数字图书馆的内容管理。
—This paper proposes an entity extraction and matching system for digital documents. Digital documents usually contain many links to their relevant information, but they do not cover all the links. Entity extraction and matching systems are used to detect such implicit links. They usually consist of several steps such as parsing, dictionary matching, and classification. Some of these steps, however, inevitably cause errors, which must be managed properly so that the process of subsequent steps is not degraded. We have therefore been developing an entity extraction and matching system focusing on managing the errors incurred at each step. This paper overviews the system and explains some techniques we have developed to improve the quality of entity extraction and matching because the system can be a key solution to content management for institutional repositories and academic societies as well as digital libraries.