Name Disambiguation Using Many-to-One features

Name Disambiguation Using Many-to-One features
复制标题

使用多对一功能消除名称歧义

DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

随着Web的飞速发展,越来越多的实体信息出现在Web上,包括他们的个人资料信息、包含他们的想法、活动、言论等的Web日志等,然而,有许多实体共享相同的名称。本文提出了一种利用多对一特征估计同名实体个数的方法。其基本思想是,实体不可能共享所有其他特征,即使它们具有相同的名称。提出了一种基于Web的关键特征提取方法,并将其联合收割机用于实体数目的估计。此外,我们还提出了一种方法来识别和过滤虚假信息,这将使我们感到困惑。网络上出现的同名指称物由于缺乏识别特征而难以区分。最初,名称是用于将一个实体与其他实体区分开来的关键特征。应该是多对一的关系,以防止误解。然而,随着越来越多的实体出现在Web上,这种关系被打破,成为多对多的关系。一个名称可能涉及数十个或数百个实体。在Web上,由于同音异义词的普遍存在,人们不能仅仅用名称来识别实体。我们的目的是估计人名表中人名的指称个数的下限,因此在一般的人名消歧过程中忽略了人名识别。我们提出了一种基于选择一些具有多对一关系的特征的方法,这些特征包括实体的轮廓和实体之间的关系。至于一个人的名字,这个人的生日和他父母的名字可能是重要的特征。利用这些特征,结合人名,我们就能够将这个人与其他同音异义词区分开来。如何选择特征和提取特征是首先要考虑的问题。Web上的数据往往是冗余的,并且充满了特征模式。我们使用迭代模式关系提取,使我们的方法可扩展。许多网页只包含实体名称,而不包含我们想要的功能。我们还将使用查询扩展来避免数据稀疏。
As the Web increase drastically, more and more entity information come to appear on Web, including their profile information, their web log containing their idea, activity, speech and so on. However, there are many entities sharing same names. Such entities include persons, locations and so on. This paper presents an approach to estimate the number of entities sharing same name by employing many-to-one features. The basic idea is that the entities are not likely to share all other features even if they have the same name. We list some strategies for selecting key features, present an approach to extract the features on Web, and combine them to estimate the entity number. What is more, we also give a method to identify the fake information which will confuse us and filter them. The referents of a same name appearing on Web are difficult to distinguish due to lack of features which can be used to identify them. Originally, the name is a key feature used to identify one entity from the others. should be a many-to-one relation to prevent misunderstanding. However, while more and more entities come to appear on Web, such relation has been broken out and it becomes to be many-to-many relation. One name may potentially refer to tens or hundreds of entities. One can't just use name to identify an entity on Web because of homonym exists so commonly. Our purpose is to estimate lower bound of referents' number of a name in name list, so we ignore name recognition in general process of name disambiguation. We present a method based on choosing some special features which have many-to-one relations, including profiles of the entities and relations between entities. As to a person name, the person's birthday and his parent's name may be important features. Use these features, combined with the person's name, we are able to distinguish this person from other homonyms. How to choose features and extract features should be considered first. The data on web is always redundant and full of feature patterns. We use iterative pattern relation extraction to make our method scalable. Many web pages just contain entity name and don't contain the features we want. And we will also use query expansion to avoid data sparse.
DOI: 10.1007/3-540-44796-2_12
发表时间: 2001-09
期刊: --
影响因子: --
作者:
David A. Smith;G. Crane
通讯作者: David A. Smith;G. Crane
DOI: 10.3115/1073083.1073092
发表时间: 2002-07
期刊: --
影响因子: --
作者:
Deepak Ravichandran;E. Hovy
通讯作者: Deepak Ravichandran;E. Hovy
DOI: 10.1007/10704656_11
发表时间: 1998-03
期刊: --
影响因子: --
作者:
Sergey Brin
通讯作者: Sergey Brin
DOI: 10.1145/1060745.1060813
发表时间: 2005-05
影响因子: 5.4
作者:
Ron Bekkerman;A. McCallum
通讯作者: Ron Bekkerman;A. McCallum
DOI: 10.3115/974557.974587
发表时间: 1997-03
期刊: --
影响因子: --
作者:
Nina Wacholder;Yael Ravin;Misook Choi
通讯作者: Nina Wacholder;Yael Ravin;Misook Choi