Shades of Gray: Seeing the Full Spectrum of Practical Data De-Identification

Shades of Gray: Seeing the Full Spectrum of Practical Data De-Identification
复制标题

灰色阴影:了解实际数据去识别的全谱

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
Kelsey Finch
Kelsey Finch
中科院分区:
--
文献类型:
--
作者:
Jules Polonetsky;Omer Tene;Kelsey Finch

文献摘要

被引文献

相似文献

隐私和数据安全领域争论最激烈的问题之一是个人数据的可识别性及其技术推论--去身份识别。去身份识别是从组织收集、存储和使用的数据中删除个人身份信息的过程。去身份识别曾被视为让组织在最大限度地减少隐私和数据安全风险的同时获得数据好处的灵丹妙药,但随着学术研究论文和流行媒体报道强调其缺点,去身份识别受到了严格的审查。与此同时,世界各地的组织必须继续依靠广泛的技术、行政和法律措施来降低个人数据的可识别性,以便能够进行关键用途和有价值的研究,同时保护个人的身份和隐私。围绕个人身份信息这一术语的轮廓的争论仍在继续,这引发了一系列法律和监管保护。科学家和监管机构经常将某些类别的信息称为“个人”信息,尽管企业和行业组织将它们定义为“非个人身份”或“非个人身份”。这场辩论的利害关系很大。虽然不是万无一失的,但去身份识别技术通过启用重要的公共和私人研究来释放价值,允许维护和使用--在某些情况下,共享和发布--有价值的信息,同时减轻隐私风险。本文提出了根据多个可识别性等级来校准数据法律规则的参数,同时还评估了其他因素,如组织的保障和控制,以及数据的敏感性、可访问性和持久性。它建立在新兴学术的基础上,该学术认为,政策制定者不应将数据视为非黑即白的二分法,而应以各种不同的灰色来看待数据;并就如何在可识别性类别之间设定重要的法律和技术边界提供指导。它敦促制定政策,鼓励各组织避免明确识别,并部署精心设计的保障和控制措施,同时保持数据集的效用。
One of the most hotly debated issues in privacy and data security is the notion of identifiability of personal data and its technological corollary, de-identification. De-identification is the process of removing personally identifiable information from data collected, stored and used by organizations. Once viewed as a silver bullet allowing organizations to reap the benefits of data while minimizing privacy and data security risks, de-identification has come under intense scrutiny with academic research papers and popular media reports highlighting its shortcomings. At the same time, organizations around the world necessarily continue to rely on a wide range of technical, administrative and legal measures to reduce the identifiability of personal data to enable critical uses and valuable research while providing protection to individuals’ identity and privacy. The debate around the contours of the term personally identifiable information, which triggers a set of legal and regulatory protections, continues to rage. Scientists and regulators frequently refer to certain categories of information as “personal” even as businesses and trade groups define them as “de-identified” or “non-personal.” The stakes in the debate are high. While not foolproof, de-identification techniques unlock value by enabling important public and private research, allowing for the maintenance and use – and, in certain cases, sharing and publication – of valuable information, while mitigating privacy risk. This paper proposes parameters for calibrating legal rules to data depending on multiple gradations of identifiability, while also assessing other factors such as an organization’s safeguards and controls, as well as the data’s sensitivity, accessibility and permanence. It builds on emerging scholarship that suggests that rather than treat data as a black or white dichotomy, policymakers should view data in various shades of gray; and provides guidance on where to place important legal and technical boundaries between categories of identifiability. It urges the development of policy that creates incentives for organizations to avoid explicit identification and deploy elaborate safeguards and controls, while at the same time maintaining the utility of data sets.