Implementing a hash-based privacy-preserving record linkage tool in the OneFlorida clinical research network

Implementing a hash-based privacy-preserving record linkage tool in the OneFlorida clinical research network
复制标题

DOI:
10.1093/jamiaopen/ooz050
复制
发表时间:
2019-12-01
期刊:
影响因子:
2.1
通讯作者:
Hogan, William
Hogan, William
中科院分区:
其他
文献类型:
--
作者:
Bian, Jiang;Loiacono, Alexander;Hogan, William

文献摘要

被引文献

相似文献

目的:在一个大型研究网络中实现一个在真实环境中执行确定性隐私保护记录链接(RL)的开源工具。材料和方法:利用公开的选民登记数据,我们学习了2条有效的确定性链接规则。然后,我们使用两个手动管理的黄金标准数据集验证了两个规则的性能,这些数据集链接了来自两个来源的电子健康记录和索赔数据。我们开发了一个基于PYTHON的开源工具OneFL DeDuper,该工具(1)使用加密的单向散列函数创建患者准标识组合的种子散列码,以实现隐私保护;(2)通过高精度和合理的召回率匹配散列码,使用中央代理链接和删除患者记录。结果:我们将OneF1 DeDuper(https://github.com/ufbmi/onefl-deduper))部署在OneFL,一个基于州的临床研究网络,作为国家以患者为中心的临床研究网络(PCORnet)的一部分。使用金标数据集,我们获得了97.25的准确率,接近99.7%,召回率为75.5%。使用该工具,我们对6个医疗保健合作伙伴和佛罗里达医疗补助计划中的170万名独特患者的350万条记录(总共1500万条)进行了重复数据删除。我们通过检查相关队列的不同疾病特征展示了RL的好处。结论:包括隐私风险考虑、政策法规、数据可用性和质量以及计算资源在内的许多因素都会影响RL解决方案在现实世界中的构建方式。然而,RL对于提高网络中的数据质量是一项重要的任务,以便我们能够从这些海量数据资源中获得可靠的科学发现。
Objective: To implement an open-source tool that performs deterministic privacy-preserving record linkage (RL) in a real-world setting within a large research network.Materials and Methods: We learned 2 efficient deterministic linkage rules using publicly available voter registration data. We then validated the 2 rules' performance with 2 manually curated gold-standard datasets linking electronic health records and claims data from 2 sources. We developed an open-source Python-based tool-OneFL Deduper-that (1) creates seeded hash codes of combinations of patients' quasi-identifiers using a cryptographic one-way hash function to achieve privacy protection and (2) links and deduplicates patient records using a central broker through matching of hash codes with a high precision and reasonable recall.Results: We deployed the OneFl Deduper (https://github.com/ufbmi/onefl-deduper) in the OneFlorida, a state-based clinical research network as part of the national Patient-Centered Clinical Research Network (PCORnet). Using the gold-standard datasets, we achieved a precision of 97.25 similar to 99.7% and a recall of 75.5%. With the tool, we deduplicated similar to 3.5 million (out of similar to 15 million) records down to 1.7 million unique patients across 6 health care partners and the Florida Medicaid program. We demonstrated the benefits of RL through examining different disease profiles of the linked cohorts.Conclusions: Many factors including privacy risk considerations, policies and regulations, data availability and quality, and computing resources, can impact how a RL solution is constructed in a real-world setting. Nevertheless, RL is a significant task in improving the data quality in a network so that we can draw reliable scientific discoveries from these massive data resources.