Mining for personal name aliases on the web

Mining for personal name aliases on the web
复制标题

DOI:
10.1145/1367497.1367679
复制
发表时间:
2008-04
期刊:
--
影响因子:
--
通讯作者:
Danushka Bollegala;Taiki Honma;Y. Matsuo;M. Ishizuka
Danushka Bollegala;Taiki Honma;Y. Matsuo;M. Ishizuka
中科院分区:
其他
文献类型:
--
作者:
Danushka Bollegala;Taiki Honma;Y. Matsuo;M. Ishizuka

文献摘要

被引文献

相似文献

我们提出了一种新的方法来找到一个给定的名字从网络上的别名。我们利用一组已知的名字和他们的别名作为训练数据和提取词汇模式,传达信息的别名的名字从文本片段返回的Web搜索引擎。然后使用这些模式来查找给定名称的候选别名。我们使用锚文本和超链接来设计一个词共现模型,并定义许多排名分数来评估一个名字和它的候选别名之间的关联。所提出的方法优于许多基线和以前的工作别名提取的数据集上的个人姓名,实现了统计上显着的平均倒数排名为0.6718。此外,使用所提出的方法提取的别名在关系检测任务中提高了20%的召回率。
We propose a novel approach to find aliases of a given name from the web. We exploit a set of known names and their aliases as training data and extract lexical patterns that convey information related to aliases of names from text snippets returned by a web search engine. The patterns are then used to find candidate aliases of a given name. We use anchor texts and hyperlinks to design a word co-occurrence model and define numerous ranking scores to evaluate the association between a name and its candidate aliases. The proposed method outperforms numerous baselines and previous work on alias extraction on a dataset of personal names, achieving a statistically significant mean reciprocal rank of 0.6718. Moreover, the aliases extracted using the proposed method improve recall by 20% in a relation-detection task.