Person Name Disambiguation in Web Pages Using Social Network, Compound Words and Latent Topics

Person Name Disambiguation in Web Pages Using Social Network, Compound Words and Latent Topics
复制标题

DOI:
10.1007/978-3-540-68125-0_24
复制
发表时间:
2008-05
期刊:
--
影响因子:
--
通讯作者:
Shingo Ono;Issei Sato;Minoru Yoshida;Hiroshi Nakagawa
Shingo Ono;Issei Sato;Minoru Yoshida;Hiroshi Nakagawa
中科院分区:
其他
文献类型:
--
作者:
Shingo Ono;Issei Sato;Minoru Yoshida;Hiroshi Nakagawa

文献摘要

相似文献

万维网(WWW)提供关于人的许多信息,并且近年来WWW搜索引擎已经普遍用于了解人。然而,许多人具有相同的姓名,并且这种模糊性通常导致一个人名的搜索结果包括关于几个不同的人的网页。我们提出了一个新的框架,有以下三个组成部分的过程中的人名消歧。通过发现命名实体的共现来提取社交网络信息,基于关键复合词的出现来测量文档相似性,基于Dirichlet过程unigram混合模型从文档中推断主题信息。使用一个实际的Web文档数据集的实验表明,我们的框架是有前途的。
The World Wide Web (WWW) provides much information about persons, and in recent years WWW search engines have been commonly used for learning about persons. However, many persons have the same name and that ambiguity typically causes the search results of one person name to include Web pages about several different persons. We propose a novel framework for person name disambiguation that has the following three components processes. Extraction of social network information by finding co-occurrences of named entities, Measurement of document similarities based on occurrences of key compound words, Inference of topic information from documents based on the Dirichlet process unigram mixture model. Experiments using an actual Web document dataset show that the result of our framework is promising.