A systematic review of re-identification attacks on health data.

A systematic review of re-identification attacks on health data.
复制标题

DOI:
10.1371/journal.pone.0028071
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Malin B
Malin B
中科院分区:
综合性期刊3区
文献类型:
--
作者:
El Emam K;Jonker E;Arbuckle L;Malin B

文献摘要

被引文献

相似文献

大多数司法管辖区的隐私立法允许在未经患者同意的情况下出于次要目的披露健康数据(如果已去识别化)。最近医学、法律和计算机科学文献中的一些文章认为,去识别方法不能提供足够的保护,因为它们很容易被逆转。如果是这种情况,它将对健康信息的披露方式产生重大影响,包括:(a) 可能限制其用于研究等次要目的的可用性,以及 (b) 导致披露更多可识别的健康信息。我们在本次系统综述中的目标是:(a) 描述对健康数据的已知重新识别攻击的特征,并将其与对其他类型数据的重新识别攻击进行对比,(b) 计算在这些攻击中已正确重新识别的记录的总体比例,以及 (c) 评估这些攻击是否表明当前去识别方法存在弱点。检索在 IEEE Xplore、ACM Digital Library 和 PubMed 中进行。经过筛选,确定了代表不同攻击的十四篇符合条件的文章。平均而言,所有研究中大约四分之一的记录被重新识别(0.26,95% CI 0.046–0.478),健康数据攻击的记录为 0.34(95% CI 0–0.744)。正如宽置信区间所证明的那样,比例存在相当大的不确定性,并且重新识别的记录的平均比例对未发表的研究很敏感。十四次攻击中有两次是使用使用现有标准去识别化的数据进行的。这些攻击中只有一次针对健康数据,成功率为 0.00013。目前的证据显示重新识别率很高,但主要是针对未根据现有标准去识别的数据进行的小规模研究。这些证据不足以得出有关去识别方法有效性的结论。
Privacy legislation in most jurisdictions allows the disclosure of health data for secondary purposes without patient consent if it is de-identified. Some recent articles in the medical, legal, and computer science literature have argued that de-identification methods do not provide sufficient protection because they are easy to reverse. Should this be the case, it would have significant and important implications on how health information is disclosed, including: (a) potentially limiting its availability for secondary purposes such as research, and (b) resulting in more identifiable health information being disclosed. Our objectives in this systematic review were to: (a) characterize known re-identification attacks on health data and contrast that to re-identification attacks on other kinds of data, (b) compute the overall proportion of records that have been correctly re-identified in these attacks, and (c) assess whether these demonstrate weaknesses in current de-identification methods. Searches were conducted in IEEE Xplore, ACM Digital Library, and PubMed. After screening, fourteen eligible articles representing distinct attacks were identified. On average, approximately a quarter of the records were re-identified across all studies (0.26 with 95% CI 0.046–0.478) and 0.34 for attacks on health data (95% CI 0–0.744). There was considerable uncertainty around the proportions as evidenced by the wide confidence intervals, and the mean proportion of records re-identified was sensitive to unpublished studies. Two of fourteen attacks were performed with data that was de-identified using existing standards. Only one of these attacks was on health data, which resulted in a success rate of 0.00013. The current evidence shows a high re-identification rate but is dominated by small-scale studies on data that was not de-identified according to existing standards. This evidence is insufficient to draw conclusions about the efficacy of de-identification methods.