Detecting, categorizing and clustering entity mentions in Chinese text

Detecting, categorizing and clustering entity mentions in Chinese text
复制标题

DOI:
10.1145/1277741.1277852
复制
发表时间:
2007-07
期刊:
--
影响因子:
--
通讯作者:
Wenjie Li;Donglei Qian;Q. Lu;C. Yuan
Wenjie Li;Donglei Qian;Q. Lu;C. Yuan
中科院分区:
其他
文献类型:
--
作者:
Wenjie Li;Donglei Qian;Q. Lu;C. Yuan

文献摘要

被引文献

相似文献

本文提出的工作是由内容提取的实际需求以及 ACE 计划的可用数据源和评估基准推动的。我们对中文实体检测和识别(EDR)任务特别感兴趣。这项任务给我们带来了一些与语言无关和语言相关的挑战,例如由于提取目标的复杂性和分词问题等而引起的。在本文中,我们提出了一种新颖的解决方案来缓解任务中的特殊问题。提及检测利用机器学习方法和基于字符的模型。它分别操纵不同类型的实体和不同的构成单元(即范围和头)。通过集成基于最具体优先和最接近优先规则的成对聚类算法,引用同一实体的提及被链接在一起。提及和实体的类型由头部驱动的分类方法确定。所实现的系统在EDR 2005中文语料库上的评估中获得了66.1的ACE值,这已经是顶级结果之一。还讨论和分析了提及检测和聚类的替代方法。
The work presented in this paper is motivated by the practical need for content extraction, and the available data source and evaluation benchmark from the ACE program. The Chinese Entity Detection and Recognition (EDR) task is of particular interest to us. This task presents us several language-independent and language-dependent challenges, e.g. rising from the complication of extraction targets and the problem of word segmentation, etc. In this paper, we propose a novel solution to alleviate the problems special in the task. Mention detection takes advantages of machine learning approaches and character-based models. It manipulates different types of entities being mentioned and different constitution units (i.e. extents and heads) separately. Mentions referring to the same entity are linked together by integrating most-specific-first and closest-first rule based pairwise clustering algorithms. Types of mentions and entities are determined by head-driven classification approaches. The implemented system achieves ACE value of 66.1 when evaluated on the EDR 2005 Chinese corpus, which has been one of the top-tier results. Alternative approaches to mention detection and clustering are also discussed and analyzed.