Evaluation of approximate comparison methods on Bloom filters for probabilistic linkage.

Evaluation of approximate comparison methods on Bloom filters for probabilistic linkage.
复制标题

DOI:
10.23889/ijpds.v4i1.1095
复制
发表时间:
2019-05-23
影响因子:
--
通讯作者:
Ferrante, A M
Ferrante, A M
中科院分区:
其他
文献类型:
--
作者:
Brown, A P;Randall, S M;Ferrante, A M

文献摘要

被引文献

相似文献

前言:数据链接中对隐私保护的需求推动了隐私保护记录链接(PPRL)技术的发展。一种流行的技术使用Bloom Filter,通过密码分析、修改和散列变化来优化隐私,一直是该领域许多研究的重点。由于Bloom Filters在概率框架中的应用很少,因此关于Bloom Filter字段之间的近似匹配是否可以提高链接质量的信息有限。目的:在本研究中,我们在记录链接的Fellegi-Sunter模型的背景下评估了三种Bloom过滤器近似比较方法的有效性:Sorensen-Dice系数、Jaccard相似度和Hamming距离。方法:使用引入误差的合成数据集来模拟具有一系列数据质量的数据集和现实世界中的大型管理健康数据集,研究估计用于将相似度得分(对于每种近似比较方法)转换为现场和数据集水平的部分权重的部分权重曲线。使用这些部分权重曲线在每个数据集上运行重复数据删除链接。这是为了比较近似比较技术与使用简单截断相似值的链接和仅使用精确匹配的链接的结果质量。结果:使用近似比较的链接比仅使用精确比较的链接产生的结果质量明显更好。特定数据集的现场级别部分权重曲线可产生最佳质量的结果。Sorensen-Dice系数和Jaccard相似性在一系列合成数据集和真实数据集上产生了最一致的结果。结论:使用Bloom Filter相似性比较概率记录链接可以产生与未加密链接的Jaro-Winkler字符串相似的链接质量结果。使用Bloom过滤器的概率链接显著受益于相似性比较的使用,即使没有针对特定数据集进行优化,部分权重曲线也会产生最佳结果。
INTRODUCTION: The need for increased privacy protection in data linkage has driven the development of privacy-preserving record linkage (PPRL) techniques. A popular technique using Bloom filters with cryptographic analyses, modifications, and hashing variations to optimise privacy has been the focus of much research in this area. With few applications of Bloom filters within a probabilistic framework, there is limited information on whether approximate matches between Bloom filtered fields can improve linkage quality.OBJECTIVES: In this study, we evaluate the effectiveness of three approximate comparison methods for Bloom filters within the context of the Fellegi-Sunter model of recording linkage: Sorensen-Dice coefficient, Jaccard similarity and Hamming distance.METHODS: Using synthetic datasets with introduced errors to simulate datasets with a range of data quality and a large real-world administrative health dataset, the research estimated partial weight curves for converting similarity scores (for each approximate comparison method) to partial weights at both field and dataset level. Deduplication linkages were run on each dataset using these partial weight curves. This was to compare the resulting quality of the approximate comparison techniques with linkages using simple cut-off similarity values and only exact matching.RESULTS: Linkages using approximate comparisons produced significantly better quality results than those using exact comparisons only. Field level partial weight curves for a specific dataset produced the best quality results. The Sorensen-Dice coefficient and Jaccard similarity produced the most consistent results across a spectrum of synthetic and real-world datasets.CONCLUSION: The use of Bloom filter similarity comparisons for probabilistic record linkage can produce linkage quality results which are comparable to Jaro-Winkler string similarities with unencrypted linkages. Probabilistic linkages using Bloom filters benefit significantly from the use of similarity comparisons, with partial weight curves producing the best results, even when not optimised for that particular dataset.