Effect of Chinese characters on machine learning for Chinese author name disambiguation: A counterfactual evaluation
Effect of Chinese characters on machine learning for Chinese author name disambiguation: A counterfactual evaluation
复制标题
DOI:
10.1177/01655515211018171
复制
发表时间:
2021-05
影响因子:
2.4
通讯作者:
Jinseok Kim;Jenna Kim;Jinmo Kim
中科院分区:
文献类型:
--
作者:
Jinseok Kim;Jenna Kim;Jinmo Kim
Chinese author names are known to be more difficult to disambiguate than other ethnic names because they tend to share surnames and forenames, thus creating many homonyms. In this study, we demonstrate how using Chinese characters can affect machine learning for author name disambiguation. For analysis, 15K author names recorded in Chinese are transliterated into English and simplified by initialising their forenames to create counterfactual scenarios, reflecting real-world indexing practices in which Chinese characters are usually unavailable. The results show that Chinese author names that are highly ambiguous in English or with initialised forenames tend to become less confusing if their Chinese characters are included in the processing. Our findings indicate that recording Chinese author names in native script can help researchers and digital libraries enhance authority control of Chinese author names that continue to increase in size in bibliographic data.