Model Reuse in Machine Learning for Author Name Disambiguation: An Exploration of Transfer Learning

Model Reuse in Machine Learning for Author Name Disambiguation: An Exploration of Transfer Learning
复制标题

DOI:
10.1109/access.2020.3031112
复制
发表时间:
2020-10
期刊:
影响因子:
3.9
通讯作者:
Jinseok Kim;Jason Owen-Smith
Jinseok Kim;Jason Owen-Smith
中科院分区:
计算机科学3区
文献类型:
--
作者:
Jinseok Kim;Jason Owen-Smith

文献摘要

相似文献

用于作者姓名消歧的机器学习通常在为特定任务创建的标记数据的训练和测试子集上进行。因此,在异构标记数据上学习的消歧模型通常不适用于不使用相同标记数据或根本不使用任何标记数据的其他目的。本文在一个新的背景下探讨了迁移学习的想法,作者姓名消歧。我们专注于缺乏标记训练数据的消歧任务使用在为其他任务生成的标记数据上训练的模型的情况。为此,两个标记的源数据集用于训练消歧模型,以应用于缺乏标记的训练数据的三个测试目标数据集。我们的研究结果表明,迁移学习可以产生与传统机器学习相似的消歧性能,其中训练和测试数据集来自相同的标记数据源。当训练源数据集与测试目标数据集具有相似的特征分布时,迁移学习的良好性能是可能的。本研究表明,通过迁移学习,丰富的消歧模型在以前的研究可以保留和重用歧义书目数据从不同的领域和数据源,激励进一步研究如何纠正源和目标数据集之间的特征分布差异,以扩大迁移学习在作者姓名消歧的应用超出本研究中探索的模型共享。
Machine learning for author name disambiguation is usually conducted on the training and test subsets of labeled data created for a specific task. As a result, disambiguation models learned on heterogeneous labeled data are often inapplicable for other purposes that either do not use the same labeled data or do not make use of any labeled data at all. This article explores the idea of transfer learning in a new context, author name disambiguation. We focus on cases where a disambiguation task lacking labeled training data uses models trained on labeled data generated for other tasks. For this purpose, two labeled source datasets are used for training of disambiguation models to be applied to three test target datasets that are deficient of labeled training data. Our results show that transfer learning can produce disambiguation performances similar to those achievable by traditional machine learning in which training and test datasets come from the same labeled data source. The good performance through transfer learning are possible when training source datasets have similar feature distributions as test target datasets. This study suggests that through transfer learning, rich disambiguation models in previous studies can be retained and reused across ambiguous bibliographic data from different fields and data sources, motivating further research on how to correct feature distribution differences between source and target datasets to expand the application of transfer learning in author name disambiguation beyond the model sharing explored in this research.