The impact of imbalanced training data on machine learning for author name disambiguation

The impact of imbalanced training data on machine learning for author name disambiguation
复制标题

DOI:
10.1007/s11192-018-2865-9
复制
发表时间:
2018-10-01
期刊:
影响因子:
3.9
通讯作者:
Kim, Jenna
Kim, Jenna
中科院分区:
管理学3区
文献类型:
--
作者:
Kim, Jinseok;Kim, Jenna

文献摘要

被引文献

相似文献

在用于作者姓名消歧的监督机器学习中,负训练数据通常显着大于正训练数据。本文研究了负训练数据与正训练数据的比率如何影响机器学习算法的性能,以消除书目记录中作者姓名的歧义。在多个标记数据集上,三个分类器(逻辑回归、朴素贝叶斯和随机森林)通过代表性特征进行训练,例如从相同训练数据中提取的合著者姓名和标题词,但具有不同的正负训练数据比率。结果表明,增加负训练数据可以提高消歧性能,但性能会提高几个百分点,有时甚至会降低性能。即使使用正负训练数据的基本比例 (1:1),逻辑贝叶斯和朴素贝叶斯也能学习最佳消歧模型。此外,随机森林的性能改进往往会在 1:10 后(与 1:15 类似)迅速饱和。这些发现意味着,与使用所有训练数据的常见做法相反,可以使用部分负训练数据来训练名称消歧算法,而不会降低太多消歧性能,同时提高计算效率。这项研究呼吁作者姓名消歧学者更多地关注从不平衡数据中进行机器学习的方法。
In supervised machine learning for author name disambiguation, negative training data are often dominantly larger than positive training data. This paper examines how the ratios of negative to positive training data can affect the performance of machine learning algorithms to disambiguate author names in bibliographic records. On multiple labeled datasets, three classifiers-Logistic Regression, Naive Bayes, and Random Forest-are trained through representative features such as coauthor names, and title words extracted from the same training data but with various positive-to-negative training data ratios. Results show that increasing negative training data can improve disambiguation performance but with a few percent of performance gains and sometimes degrade it. Logistic and Naive Bayes learn optimal disambiguation models even with a base ratio (1:1) of positive and negative training data. Also, the performance improvement by Random Forest tends to quickly saturate roughly after 1:10 similar to 1:15. These findings imply that contrary to the common practice using all training data, name disambiguation algorithms can be trained using part of negative training data without degrading much disambiguation performance while increasing computational efficiency. This study calls for more attention from author name disambiguation scholars to methods for machine learning from imbalanced data.