Name Disambiguation Using Many-to-One features
Name Disambiguation Using Many-to-One features
复制标题
使用多对一功能消除名称歧义
DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
中科院分区:
文献类型:
--
作者:
As the Web increase drastically, more and more entity information come to appear on Web, including their profile information, their web log containing their idea, activity, speech and so on. However, there are many entities sharing same names. Such entities include persons, locations and so on. This paper presents an approach to estimate the number of entities sharing same name by employing many-to-one features. The basic idea is that the entities are not likely to share all other features even if they have the same name. We list some strategies for selecting key features, present an approach to extract the features on Web, and combine them to estimate the entity number. What is more, we also give a method to identify the fake information which will confuse us and filter them. The referents of a same name appearing on Web are difficult to distinguish due to lack of features which can be used to identify them. Originally, the name is a key feature used to identify one entity from the others. should be a many-to-one relation to prevent misunderstanding. However, while more and more entities come to appear on Web, such relation has been broken out and it becomes to be many-to-many relation. One name may potentially refer to tens or hundreds of entities. One can't just use name to identify an entity on Web because of homonym exists so commonly. Our purpose is to estimate lower bound of referents' number of a name in name list, so we ignore name recognition in general process of name disambiguation. We present a method based on choosing some special features which have many-to-one relations, including profiles of the entities and relations between entities. As to a person name, the person's birthday and his parent's name may be important features. Use these features, combined with the person's name, we are able to distinguish this person from other homonyms. How to choose features and extract features should be considered first. The data on web is always redundant and full of feature patterns. We use iterative pattern relation extraction to make our method scalable. Many web pages just contain entity name and don't contain the features we want. And we will also use query expansion to avoid data sparse.
登录
查看更多内容
DOI:
10.1007/3-540-44796-2_12
发表时间:
2001-09
期刊:
--
影响因子:
--
作者:
David A. Smith;G. Crane
通讯作者:
David A. Smith;G. Crane
DOI:
10.3115/1073083.1073092
发表时间:
2002-07
期刊:
--
影响因子:
--
作者:
Deepak Ravichandran;E. Hovy
通讯作者:
Deepak Ravichandran;E. Hovy
DOI:
10.1007/10704656_11
发表时间:
1998-03
期刊:
--
影响因子:
--
作者:
Sergey Brin
通讯作者:
Sergey Brin
影响因子:
5.4
作者:
Ron Bekkerman;A. McCallum
通讯作者:
Ron Bekkerman;A. McCallum
DOI:
10.3115/974557.974587
发表时间:
1997-03
期刊:
--
影响因子:
--
作者:
Nina Wacholder;Yael Ravin;Misook Choi
通讯作者:
Nina Wacholder;Yael Ravin;Misook Choi