Identification of Personal Name Aliases on the Web

Identification of Personal Name Aliases on the Web
复制标题

DOI:
--
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
Danushka Bollegala;Taiki Honma;Y. Matsuo;M. Ishizuka
Danushka Bollegala;Taiki Honma;Y. Matsuo;M. Ishizuka
中科院分区:
其他
文献类型:
--
作者:
Danushka Bollegala;Taiki Honma;Y. Matsuo;M. Ishizuka

文献摘要

相似文献

提取实体的别名对于诸如实体之间的关系的识别、web搜索和实体消歧等各种任务都是重要的。为了正确地提取实体之间的关系,必须首先识别这些实体。我们提出了一种新的方法来flnd别名的一个给定的名字使用自动提取的词汇模式。我们利用一组已知的名字和他们的别名作为训练数据和提取词汇模式,传达信息的别名的名字从文本片段返回的Web搜索引擎。然后使用这些模式来查找给定名称的候选别名。我们使用锚文本来设计一个词共现模型,并使用它来定义各种排名分数,以衡量一个名字和一个候选别名之间的关联。使用支持向量机将排名分数与基于页面计数的关联措施相结合,以利用强大的别名检测方法。所提出的方法优于许多基线和以前的工作上的别名提取的数据集上的个人姓名,实现了统计显着的平均倒数排名0:6718。使用位置名称和日本人的名字的数据集进行的实验表明,扩展所提出的方法来提取不同类型的命名实体和其他语言的别名的可能性。此外,使用所提出的方法提取的别名在关系检测任务中提高了20%的召回率。
Extracting aliases of an entity is important for various tasks such as identiflcation of relations among entities, web search and entity disambiguation. To extract relations among entities properly, one must flrst identify those entities. We propose a novel approach to flnd aliases of a given name using automatically extracted lexical patterns. We exploit a set of known names and their aliases as training data and extract lexical patterns that convey information related to aliases of names from text snippets returned by a web search engine. The patterns are then used to flnd candidate aliases of a given name. We use anchor texts to design a word cooccurrence model and use it to deflne various ranking scores to measure the association between a name and a candidate alias. The ranking scores are integrated with page-countbased association measures using support vector machines to leverage a robust alias detection method. The proposed method outperforms numerous baselines and previous work on alias extraction on a dataset of personal names, achieving a statistically signiflcant mean reciprocal rank of 0:6718. Experiments carried out using a dataset of location names and Japanese personal names suggest the possibility of extending the proposed method to extract aliases for difierent types of named entities and for other languages. Moreover, the aliases extracted using the proposed method improve recall by 20% in a relation-detection task.