Saying it’s Anonymous Doesn't Make It So: Re-identifications of “anonymized” law school data
Saying it’s Anonymous Doesn't Make It So: Re-identifications of “anonymized” law school data
复制标题
说是匿名并不代表事实如此:“匿名”法学院数据的重新识别
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
A. Perry
中科院分区:
文献类型:
--
作者:
L. Sweeney;Michael von Loewenfeldt;A. Perry
Society trusts data privacy practitioners to make decisions about which fields of personal income, medical, or educational information can be shared publicly in accordance with laws and standards. How good are the decisions they make? They don’t have to publish the protocols they use, and they often prohibit others from telling them about vulnerabilities found in the data. So, in the silence, these practitioners circularly assert that there are no problems. We had a unique opportunity in a legal setting to examine the real-world decisionmaking of a team of accomplished data privacy experts and to test the quality and accuracy of the decisions they make. The litigation, Richard Sander et. al v. State Bar of California et. al., was over whether the release of requested data was required by California law [1]. During the lawsuit, an expert team of data privacy practitioners proposed four “best practice” protocols that they asserted were sufficient to protect the privacy of individuals whose Sweeney L, Von Loewenfeldt M, Perry M. Saying it’s Anonymous Doesn’t Make It So: Re-identifications of “anonymized” law school data. Technology Science. 2018111301. November 13, 2018. http://techscience.org/a/2018111301 3 information was in the data. All four protocols claimed to leverage approaches widely used today in government, corporate, and research practice. This paper presents their protocols and shows, based on analysis that was made public during the trial, vulnerabilities that each protocol had to re-identifications – the ability to associate real names to “anonymized” data records. Results summary: One protocol used a physical data enclave, two purported to produce a kanonymous version of the data, and a fourth protocol developed a statistical model of the data. None of the protocols provided the privacy protection promised or commensurate with common expectations under public records laws. We demonstrate important lessons: (1) kanonymity guarantees that an adversary cannot do better than guessing that a name matches to at least k records or, vice versa, that at least k people ambiguously match to a record. None of the “k-anonymity” protocols were actually k-anonymous. (2) In today’s datarich, networked society, the k constraint must be enforced across all fields or scientific justification provided to exclude a field. The “k-anonymity” protocols excluded some fields from k protection void of analytical rationale. We demonstrated ways to use those fields to help put names to records de-identified by these protocols. (3) We found small group reidentifications in all their protocols that were as harmful as unique re-identifications. (4) The physical data enclave limited access to the data, but still could not thwart hiding or memorizing sensitive information on targeted individuals. (5) All four protocols left the records of Black and Hispanic test-takers significantly more identifiable than the records of Whites. The Superior Court of California denied Sander’s request for compelled disclosure of the data, and the California Court of Appeals upheld the decision. Our findings demonstrate how adversarial testing on de-identified data can point out vulnerabilities and improve realworld practice.
DOI:
--
发表时间:
2017
期刊:
Technology science
影响因子:
--
作者:
Sweeney,Latanya;Yoo,JiSu;Perovich,Laura;Boronow,KatherineE;Brown,Phil;Brody,JuliaGreen
通讯作者:
Brody,JuliaGreen