Citation data clustering for author name disambiguation

Citation data clustering for author name disambiguation
复制标题

DOI:
10.1145/1366804.1366884
复制
发表时间:
2007-06
期刊:
--
影响因子:
--
通讯作者:
Tomonari Masada;A. Takasu;J. Adachi
Tomonari Masada;A. Takasu;J. Adachi
中科院分区:
其他
文献类型:
--
作者:
Tomonari Masada;A. Takasu;J. Adachi

文献摘要

相似文献

本文提出了一种基于引文数据聚类的作者姓名消歧方法。在科学论文的参考文献部分出现的大多数引文数据都包括合作者的名字和他们的首字母缩写。因此,我们经常使用这样的缩写名称来搜索引文数据,例如:“S. Lee”或“J. Chen”,并因此在搜索结果中获得许多不相关的数据,因为这样的缩写名称指的是许多不同的人。在本文中,我们提出了一种引文数据聚类的方法,该方法构建的每一个聚类只包含一个唯一作者对应的引文数据。我们的聚类方法是基于一个概率模型,它是朴素贝叶斯混合模型的扩展。由于我们的模型有两个隐变量,我们称之为双变量混合模型。在评价实验中,我们使用了著名的DBLP数据集。结果表明,与朴素贝叶斯混合模型相比,双变量混合模型在查准率和查全率之间取得了更好的平衡。
In this paper, we propose a new method of citation data clustering for author name disambiguation. Most citation data appearing in the reference section of scientific papers include the coauthor first names with their initials. Hence, we often search citation data by using such an abbreviated name, e.g. "S. Lee" or "J. Chen", and consequently obtain many irrelevant data in the search result, because such an abbreviated name refers to many different persons. In this paper, we propose a method of citation data clustering to construct clusters each of which includes only citation data corresponding to a unique author. Our clustering method is based on a probabilistic model which is an extension of the naive Bayes mixture model. Since our model has two hidden variables, we call it two-variable mixture model. In the evaluation experiment, we used the well-known DBLP data set. The results show that the two-variable mixture model can achieve a better balance between precision and recall than the naive Bayes mixture model.