Saying it’s Anonymous Doesn't Make It So: Re-identifications of “anonymized” law school data

Saying it’s Anonymous Doesn't Make It So: Re-identifications of “anonymized” law school data
复制标题

说是匿名并不代表事实如此:“匿名”法学院数据的重新识别

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
A. Perry
A. Perry
中科院分区:
--
文献类型:
--
作者:
L. Sweeney;Michael von Loewenfeldt;A. Perry

文献摘要

参考文献

被引文献

相似文献

社会信任数据隐私从业者根据法律和标准决定个人收入、医疗或教育信息的哪些领域可以公开共享。他们所做的决定有多好?他们不必公布他们使用的协议,而且他们经常禁止其他人告诉他们在数据中发现的漏洞。因此,在沉默中,这些实践者循环地断言没有问题。在法律环境中,我们有一个独特的机会来检查一个经验丰富的数据隐私专家团队在现实世界中的决策过程,并测试他们决策的质量和准确性。诉讼,理查德·桑德等人。Al诉加利福尼亚州州律师等人案。关于加州法律是否要求公布所要求的数据[1]。在诉讼中,一个由数据隐私从业人员组成的专家团队提出了四个“最佳实践”协议,他们声称这些协议足以保护斯威尼·L、冯·洛文菲尔特·M、佩里·M所在的个人的隐私。他们说,匿名者没有做到这一点:重新识别“匿名”的法学院数据。科技科学。2018111301。2018年11月13日。数据中有http://techscience.org/a/2018111301 3的信息。所有四个协议都声称利用了当今在政府、企业和研究实践中广泛使用的方法。本文介绍了他们的协议,并根据在试验期间公布的分析,展示了每个协议必须重新识别的漏洞-将真实姓名与“匿名”数据记录相关联的能力。结果摘要:一项协议使用了实物数据飞地,两项协议声称产生了数据的匿名版本,第四项协议开发了数据的统计模型。这些议定书都没有提供承诺的隐私保护,也没有与公共记录法规定的共同期望相称的隐私保护。我们展示了重要的教训:(1)匿名性保证对手只能猜测一个名字与至少k条记录匹配,反之亦然,至少k人与一条记录模棱两可地匹配。没有一种“k-匿名”协议实际上是k-匿名的。(2)在当今数据丰富、网络化的社会中,必须在所有领域强制实施k约束,或者必须提供科学理由来排除某个领域。“k-匿名性”协议将一些领域排除在缺乏分析理由的k保护之外。我们演示了如何使用这些字段帮助将名称添加到由这些协议取消标识的记录中。(3)在所有方案中,我们发现小群体再鉴定与唯一再鉴定一样有害。(4)物理数据飞地限制了对数据的访问,但仍无法阻止隐藏或记忆目标个人的敏感信息。(5)所有四个协议都使黑人和西班牙裔考生的记录比白人的记录更容易识别。加州高等法院驳回了桑德关于强制披露数据的请求,加州上诉法院维持了这一决定。我们的发现表明,对未识别的数据进行对抗性测试可以指出漏洞并改进现实世界的实践。
Society trusts data privacy practitioners to make decisions about which fields of personal income, medical, or educational information can be shared publicly in accordance with laws and standards. How good are the decisions they make? They don’t have to publish the protocols they use, and they often prohibit others from telling them about vulnerabilities found in the data. So, in the silence, these practitioners circularly assert that there are no problems. We had a unique opportunity in a legal setting to examine the real-world decisionmaking of a team of accomplished data privacy experts and to test the quality and accuracy of the decisions they make. The litigation, Richard Sander et. al v. State Bar of California et. al., was over whether the release of requested data was required by California law [1]. During the lawsuit, an expert team of data privacy practitioners proposed four “best practice” protocols that they asserted were sufficient to protect the privacy of individuals whose Sweeney L, Von Loewenfeldt M, Perry M. Saying it’s Anonymous Doesn’t Make It So: Re-identifications of “anonymized” law school data. Technology Science. 2018111301. November 13, 2018. http://techscience.org/a/2018111301 3 information was in the data. All four protocols claimed to leverage approaches widely used today in government, corporate, and research practice. This paper presents their protocols and shows, based on analysis that was made public during the trial, vulnerabilities that each protocol had to re-identifications – the ability to associate real names to “anonymized” data records. Results summary: One protocol used a physical data enclave, two purported to produce a kanonymous version of the data, and a fourth protocol developed a statistical model of the data. None of the protocols provided the privacy protection promised or commensurate with common expectations under public records laws. We demonstrate important lessons: (1) kanonymity guarantees that an adversary cannot do better than guessing that a name matches to at least k records or, vice versa, that at least k people ambiguously match to a record. None of the “k-anonymity” protocols were actually k-anonymous. (2) In today’s datarich, networked society, the k constraint must be enforced across all fields or scientific justification provided to exclude a field. The “k-anonymity” protocols excluded some fields from k protection void of analytical rationale. We demonstrated ways to use those fields to help put names to records de-identified by these protocols. (3) We found small group reidentifications in all their protocols that were as harmful as unique re-identifications. (4) The physical data enclave limited access to the data, but still could not thwart hiding or memorizing sensitive information on targeted individuals. (5) All four protocols left the records of Black and Hispanic test-takers significantly more identifiable than the records of Whites. The Superior Court of California denied Sander’s request for compelled disclosure of the data, and the California Court of Appeals upheld the decision. Our findings demonstrate how adversarial testing on de-identified data can point out vulnerabilities and improve realworld practice.
HIPAA 安全港数据中的重新识别风险:一项环境健康研究数据的研究。
DOI: --
发表时间: 2017
期刊: Technology science
影响因子: --
作者:
Sweeney,Latanya;Yoo,JiSu;Perovich,Laura;Boronow,KatherineE;Brown,Phil;Brody,JuliaGreen
通讯作者: Brody,JuliaGreen