Effect of Chinese characters on machine learning for Chinese author name disambiguation: A counterfactual evaluation

Effect of Chinese characters on machine learning for Chinese author name disambiguation: A counterfactual evaluation
复制标题

DOI:
10.1177/01655515211018171
复制
发表时间:
2021-05
影响因子:
2.4
通讯作者:
Jinseok Kim;Jenna Kim;Jinmo Kim
Jinseok Kim;Jenna Kim;Jinmo Kim
中科院分区:
计算机科学3区
文献类型:
--
作者:
Jinseok Kim;Jenna Kim;Jinmo Kim

文献摘要

相似文献

众所周知,中国作者的名字比其他民族的名字更难消除歧义,因为他们往往共用姓氏和名字,从而产生了许多同音异义词。在这项研究中,我们演示了使用汉字如何影响作者姓名消歧的机器学习。为了进行分析,用中文记录的 15K 个作者姓名被音译为英文,并通过初始化他们的名字来简化,以创建反事实场景,反映了现实世界中汉字通常不可用的索引实践。结果表明,对于英文中高度模糊的中文作者姓名或名字带有首字母缩写的中文作者姓名,如果将其汉字包含在处理中,往往会变得不那么混乱。我们的研究结果表明,以母文字记录中文作者姓名可以帮助研究人员和数字图书馆加强对书目数据中不断增加的中文作者姓名的权限控制。
Chinese author names are known to be more difficult to disambiguate than other ethnic names because they tend to share surnames and forenames, thus creating many homonyms. In this study, we demonstrate how using Chinese characters can affect machine learning for author name disambiguation. For analysis, 15K author names recorded in Chinese are transliterated into English and simplified by initialising their forenames to create counterfactual scenarios, reflecting real-world indexing practices in which Chinese characters are usually unavailable. The results show that Chinese author names that are highly ambiguous in English or with initialised forenames tend to become less confusing if their Chinese characters are included in the processing. Our findings indicate that recording Chinese author names in native script can help researchers and digital libraries enhance authority control of Chinese author names that continue to increase in size in bibliographic data.