A software package for the application of probabilistic anonymisation to sensitive individual-level data: a proof of principle with an example from the ALSPAC birth cohort study

A software package for the application of probabilistic anonymisation to sensitive individual-level data: a proof of principle with an example from the ALSPAC birth cohort study
复制标题

DOI:
10.14301/llcs.v9i4.478
复制
发表时间:
2018-10-01
影响因子:
0.9
通讯作者:
Burton, Paul
Burton, Paul
中科院分区:
法学4区
文献类型:
--
作者:
Avraam, Demetris;Boyd, Andy;Burton, Paul

文献摘要

被引文献

相似文献

个人级别的数据需要保护,以防止未经授权的访问,以保护敏感信息的机密性和安全性。披露风险透过隐私风险评估进行评估,并于分享及整合资料前予以控制或减至最低。从“微型数据实验室”传统(即在受控的物理位置访问)到“开放数据”(即共享个人层面的数据)的演变推动了高效匿名方法和保护控制的发展。有效的匿名化技术应增加重新识别的不确定性,同时保留数据效用,允许进行信息丰富的数据分析。“概率匿名化”就是这样一种技术,它通过添加随机噪声来改变数据。在本文中,我们描述了一个概率匿名化技术到R编写的操作软件中的实现,并通过应用于ALSPAC队列研究的哮喘相关数据分析来证明其适用性。该软件旨在供数据管理人员和用户使用,而不需要先进的统计知识。
Individual-level data require protection from unauthorised access to safeguard confidentiality and security of sensitive information. Risks of disclosure are evaluated through privacy risk assessments and are controlled or minimised before data sharing and integration. The evolution from 'Micro Data Laboratory' traditions (i.e. access in controlled physical locations) to 'Open Data' (i.e. sharing individual-level data) drives the development of efficient anonymisation methods and protection controls. Effective anonymisation techniques should increase the uncertainty surrounding re-identification while retaining data utility, allowing informative data analysis. 'Probabilistic anonymisation' is one such technique, which alters the data by addition of random noise. In this paper, we describe the implementation of one probabilistic anonymisation technique into an operational software written in R and we demonstrate its applicability through application to analysis of asthma-related data from the ALSPAC cohort study. The software is designed to be used by data managers and users without the requirement of advanced statistical knowledge.