A Heuristic-based Hierarchical Clustering Method for Author Name Disambiguation in Digital Libraries

A Heuristic-based Hierarchical Clustering Method for Author Name Disambiguation in Digital Libraries
复制标题

数字图书馆中基于启发式的作者姓名消歧层次聚类方法

DOI:
--
复制
发表时间:
2007
期刊:
Brazilian Symposium on Databases
影响因子:
--
通讯作者:
Alberto H. F. Laender
Alberto H. F. Laender
中科院分区:
--
文献类型:
--
作者:
Ricardo G. Cota;Marcos André Gonçalves;Alberto H. F. Laender

文献摘要

参考文献

被引文献

相似文献

在本文中,我们提出了一个基于层次聚类(HHC)的方法来处理的名称消歧问题。该方法基于对引用的组成部分的若干启发式和相似性度量(例如,合著者、作品名称、出版地点)。在每个阶段中,融合簇的信息被聚合,为下一轮融合提供更多的信息。使用从DBLP集合中获取的数据集的实验显示,与不考虑分层聚类的先前方法相比,增益高达12%,与监督基线相比,增益高达21%(即,SVM)和15.5%对一个无监督的(即,K-Means),使用所考虑的相同证据。
In this paper, we propose a heuristic-based hierarchical clustering (HHC) method to deal with the name disambiguation problem. The method successively fuses clusters of citations of compatible authors based on several heuristic and similarity measures on the components of the citations (e.g., co- authors, title of the work, publication venue). In each phase, the information of fused clusters is aggregated providing more information for the next round of fusion. Experiments with a dataset taken from the DBLP collection show gains up to 12% against a previous method that did not consider hierarchical clustering and up to 21% against a supervised baseline (i.e., SVM) and 15.5% against an unsupervised one (i.e., K-Means) which use the same evidence considered.
DOI: --
发表时间: --
期刊: Advanced Materials
影响因子: 29.4
作者:
F. Azuaje
通讯作者: F. Azuaje