Evaluating author name disambiguation for digital libraries: a case of DBLP

Evaluating author name disambiguation for digital libraries: a case of DBLP
复制标题

DOI:
10.1007/s11192-018-2824-5
复制
发表时间:
2018-09-01
期刊:
影响因子:
3.9
通讯作者:
Kim, Jinseok
Kim, Jinseok
中科院分区:
管理学3区
文献类型:
--
作者:
Kim, Jinseok

文献摘要

被引文献

相似文献

数字图书馆中的著者姓名歧义现象会影响到图书馆著者身份数据挖掘的研究结果。本研究评估作者姓名消歧在DBLP,一个广泛使用,但没有充分评估其消歧性能的数字图书馆。在这样做时,本研究采取了三角测量的方法,作者姓名消歧的数字图书馆,可以更好地评估其性能时,多个标记的数据集与基线比较。测试三种类型的标记数据包含5000至6 M消歧的名称,DBLP分配作者姓名相当准确地不同的作者,导致成对精度,召回率和F1的措施约0.90或以上的整体。DBLP的作者姓名消歧即使在大的歧义姓名块上也表现良好,但在区分具有相同姓名的作者方面表现不足。与其他消歧算法相比,DBLP的消歧性能是相当有竞争力的,可能是由于它的混合消歧方法相结合的算法消歧和手动纠错。讨论如下的优点和缺点,在这项研究中使用的标记数据集的未来努力,以评估作者姓名消歧的数字图书馆规模。
Author name ambiguity in a digital library may affect the findings of research that mines authorship data of the library. This study evaluates author name disambiguation in DBLP, a widely used but insufficiently evaluated digital library for its disambiguation performance. In doing so, this study takes a triangulation approach that author name disambiguation for a digital library can be better evaluated when its performance is assessed on multiple labeled datasets with comparison to baselines. Tested on three types of labeled data containing 5000 to 6 M disambiguated names, DBLP is shown to assign author names quite accurately to distinct authors, resulting in pairwise precision, recall, and F1 measures around 0.90 or above overall. DBLP's author name disambiguation performs well even on large ambiguous name blocks but deficiently on distinguishing authors with the same names. Compared to other disambiguation algorithms, DBLP's disambiguation performance is quite competitive, possibly due to its hybrid disambiguation approach combining algorithmic disambiguation and manual error correction. A discussion follows on strengths and weaknesses of labeled datasets used in this study for future efforts to evaluate author name disambiguation on a digital library scale.