Multilingual author matching across different academic databases: a case study on KAKEN, DBLP, and PubMed

Multilingual author matching across different academic databases: a case study on KAKEN, DBLP, and PubMed
复制标题

DOI:
10.1007/s11192-020-03861-3
复制
发表时间:
2021-02
期刊:
影响因子:
3.9
通讯作者:
Yuto Chikazawa;Marie Katsurai;I. Ohmukai
Yuto Chikazawa;Marie Katsurai;I. Ohmukai
中科院分区:
管理学3区
文献类型:
--
作者:
Yuto Chikazawa;Marie Katsurai;I. Ohmukai

文献摘要

被引文献

相似文献

研究人员经常使用他们的母语来表达和交流思想。要建立一个完整的个人资料,他们的英语和非英语学术出版物的列表必须建立。本文提出了一种实用的方法,跨不同的学术数据库的多语种作者匹配。我们的方法自动链接的学术记录的目标数据库的源数据库的研究人员标识符。首先,我们在目标数据库中提取了一组完整的记录,这些记录的作者姓名与源数据库中的研究人员姓名相同。然后,我们计算了多个作者相似性度量,这些度量可以在来自不同语言数据库的某些实体对中采用。最后,我们汇总了这些指标,以输出一个改进的分数,该分数表明每条记录作为研究人员工作的可能性。我们的方法被发现是很容易实现的,其性能进行了评估,在真实的数据库管理设置。实验使用DBLP和PubMed作为目标英文数据库进行。作为日本数据库,卡肯是识别研究人员信息的来源。结果表明,每个相似性度量的性能,从我们观察到,得分聚合实现了稳定的性能。我们的方法可以减少人类将各种学术贡献联系起来的努力。
Researchers often use their native languages to present and exchange ideas. To construct an individual author’s complete profile, a list of their English and non-English academic publications must be constructed. This paper presents a practical approach for multilingual author matching across different academic databases. Our approach automatically links the academic records of a target database to a researcher identifier of a source database. First, we extracted a comprehensive set of records in the target database, whose author names were identical to the researcher names in the source database. Then, we calculated multiple author similarity measures, which can be adopted in certain entity pairs from different language databases. Finally, we aggregated the measures to output an improved score that indicates the likelihood of each record as being the researcher’s work. Our method was found to be easy to implement, and its performance was evaluated in real database management settings. Experiments were conducted using DBLP and PubMed as the target English databases. As the Japanese database, KAKEN was the source for identifying researcher information. The results demonstrated each similarity measure’s performance, from which we observed that the score aggregation achieved stable performance. Our method can lessen human efforts to associate various scholarly contributions.