A Globally Optimal k-Anonymity Method for the De-Identification of Health Data

A Globally Optimal k-Anonymity Method for the De-Identification of Health Data
复制标题

DOI:
10.1197/jamia.m3144
复制
发表时间:
2009-09-01
影响因子:
6.4
通讯作者:
Bottomley, Jim
Bottomley, Jim
中科院分区:
管理学2区
文献类型:
--
作者:
El Emam, Khaled;Dankar, Fida Kamal;Bottomley, Jim

文献摘要

被引文献

相似文献

背景:隐私法中明确的患者同意要求可能会对健康研究产生负面影响,导致选择偏差和招募减少。如果收集或披露的信息被去识别化,通常会免除获得同意的立法要求。 目标:作者开发并实证评估了一种新的全局最佳去识别化算法,该算法满足 k-匿名标准,并且适合健康数据集。 设计:作者在六个公共、医院和登记数据集上,根据经验将 OLA(最佳格子匿名化)与三种现有的 k-匿名算法 Datafly、Samarati 和 Incognito 进行比较。 k 值和抑制限。测量:使用三个信息损失度量进行比较:精度、可辨别性度量和非均匀熵。还评估了每种算法的性能速度。结果:Datafly 和 Samarati 算法的信息丢失率高于 OLA 和 Incognito; OLA 在寻找全局最优的去标识化解决方案方面始终比 Incognito 更快。结论:对于健康数据集的去标识化,OLA 在信息丢失和性能方面是对现有 k-匿名算法的改进。
Background: Explicit patient consent requirements in privacy laws can have a negative impact on health research, leading to selection bias and reduced recruitment. Often legislative requirements to obtain consent are waived if the information collected or disclosed is de-identified.Objective: The authors developed and empirically evaluated a new globally optimal de-identification algorithm that satisfies the k-anonymity criterion and that is suitable for health datasets.Design: Authors compared OLA (Optimal Lattice Anonymization) empirically to three existing k-anonymity algorithms, Datafly, Samarati, and Incognito, on six public, hospital, and registry datasets for different values of k and suppression limits.Measurement: Three information loss metrics were used for the comparison: precision, discernability metric, and non-uniform entropy. Each algorithm's performance speed was also evaluated.Results: The Datafly and Samarati algorithms had higher information loss than OLA and Incognito; OLA was consistently faster than Incognito in finding the globally optimal de-identification solution.Conclusions: For the de-identification of health datasets, OLA is an improvement on existing k-anonymity algorithms in terms of information loss and performance.