Traces through Time: a Case-study of Applying Statistical Methods to Refine Algorithms for Linking Biographical Data

Traces through Time: a Case-study of Applying Statistical Methods to Refine Algorithms for Linking Biographical Data
复制标题

DOI:
--
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
M. Bell;Sonia Ranade
M. Bell;Sonia Ranade
中科院分区:
其他
文献类型:
--
作者:
M. Bell;Sonia Ranade

文献摘要

相似文献

2015年在英国国家档案馆运行的“时间痕迹”项目开发了算法和工具,将出现在历史记录中的人联系起来,并为所建立的联系分配可靠的信心措施。该方法适用于数字人文学科,包括传记研究。模糊匹配取决于是否有关于人口的背景统计数据、数据值的分布、数据质量以及错误的类型和频率。本文介绍了工作,以完善原始算法,通过实施的学习方法,其中的见解所产生的一个分析反馈到算法,以改善基线统计数据,为后续分析。我们发现,这种迭代的方法提供了显着的改进,“原始”的评分机制。它使我们能够仔细地确定要应用的模糊匹配的类型和程度,并且可以帮助平衡由于允许增加“模糊”而导致的差的精确度和由于更限制性的方法而导致的差的召回率。今后的工作将把这一方法扩展到姓名和出生日期之外,并将这些改进嵌入“时间追踪”框架和工具。
The Traces through Time project, which ran at The UK National Archives in 2015, developed algorithms and tools to link people appearing in historical records and to assign robust measures of confidence to the connections that are made. The method has application across the digital humanities, including for biographical research. Fuzzy matching relies on the availability of background statistics on the population, the distribution of data values, data quality and the type and frequency of errors. This paper describes work to refine the original algorithms through implementation of a learning approach in which insights arising from one analysis are fed back into the algorithm to improve the baseline statistics for subsequent analyses. We find that this iterative approach delivers significant improvements over 'raw’ scoring mechanisms. It enables us to carefully target the type and degree of fuzzy matching to be applied and can help balance the poor precision that results from allowing increased ‘fuzziness’ against the poor recall that arises from a more restrictive approach. Future work will extend the approach beyond names and dates of birth, and will embed these enhancements into the Traces through Time framework and tools.