A comparison of machine learning classifiers for use on historical record linkage

A comparison of machine learning classifiers for use on historical record linkage
复制标题

用于历史记录链接的机器学习分类器的比较

DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
P. Kaur
P. Kaur
中科院分区:
--
文献类型:
--
作者:
P. Kaur

文献摘要

被引文献

相似文献

用于历史记录链接的机器学习分类器的比较Pavneet Kaur,圭尔夫大学,2020年顾问:Luiza Antonie博士记录链接是在没有唯一标识符的情况下识别一个或多个数据源中相同实体的过程。通过连接两个或多个历史来源构建的纵向数据可以为我们提供关于人口随时间变化特征的有价值的信息。然而,由于没有个人身份识别资料,这种纵向数据的构建受到挑战。在这篇论文中,我们使用不同的方法和条件连接1871年和1881年的加拿大人口普查。利用支持向量机和随机森林分类器建立了记录链接系统。这些不同的方法的性能进行了比较,调查的上限实现的联动率和条件,为我们提供了该速率进行检查。实验结果表明,本文所研究的随机森林分类方法在保持不超过5%的误报率的同时,将链接率提高了3.6%。
A COMPARISON OF MACHINE LEARNING CLASSIFIERS FOR USE ON HISTORICAL RECORD LINKAGE Pavneet Kaur, May University of Guelph, 2020 Advisor: Dr. Luiza Antonie Record Linkage is the process of identifying the same entities in one or more data sources in the absence of unique identifiers. Longitudinal data constructed by linking two or more historical sources can provide us with valuable information about the characteristics of population change over time. However, the construction of such longitudinal data is challenged by the unavailability of personal identifiers. In this thesis, we link the Canadian censuses 1871 and 1881 using different methods and conditions. The Support Vector Machine and the Random Forest Classifiers are used to establish the record linkage system. The performance of these different methods is compared to investigate the upper bound achieved in the linkage rate and the conditions which provided us with that rate are inspected. Experiments show that the Random Forest classification explored in this thesis improves upon the linkage rate by 3.6% while maintaining a false positive rate no greater than 5%.