A method for managing re-identification risk from small geographic areas in Canada

A method for managing re-identification risk from small geographic areas in Canada
复制标题

DOI:
10.1186/1472-6947-10-18
复制
发表时间:
2010-04-02
影响因子:
3.5
通讯作者:
Roffey, Tyson
Roffey, Tyson
中科院分区:
医学3区
文献类型:
--
作者:
El Emam, Khaled;Brown, Ann;Roffey, Tyson

文献摘要

被引文献

相似文献

工作背景:健康数据集的一种常见披露控制做法是识别小的地理区域,并隐藏这些小区域的记录或将其聚合到更大的区域中。最近的一项研究提供了一种方法,用于根据唯一性标准确定区域是否太小。独特性标准规定,当相关变量(准标识符)上独特个体的比例接近零时,区域不再太小。然而,使用唯一性值为零是一个非常严格的阈值,并且仅适用于数据泄露的风险非常高的情况。其他的独特性阈值,已提出的健康数据为5%和20%.Methods:我们估计的独特性城市前分拣区(FSA)使用2001年加拿大人口普查数据的20%。然后,我们构建了两个逻辑回归模型来预测独特性何时大于5%和20%阈值,并使用10倍交叉验证来验证其预测准确性。结果:所有模型参数均具有显著性意义,模型的预测准确度均在0.9以上,5%和20%阈值模型的敏感性分别为0.87和0.74。通过对安大略新生儿登记处和急诊科数据集的分析,说明了模型的应用。在较高的阈值下,与0%阈值相比,被认为属于小范围的记录要少得多,因此需要采取披露控制行动。我们还为数据保管人提供了具体指导,以决定使用三个唯一性阈值中的哪一个(0%、5%、20%),具体取决于数据接收方已采取的缓解控制措施、数据披露后可能侵犯隐私的情况,以及数据接收方重新识别数据的动机和能力。我们开发的模型可用于管理小地理区域的重新识别风险。由于能够在三个可能的阈值中进行选择,数据保管人可以根据数据和接收者的性质调整“小地理区域”的定义。
Background: A common disclosure control practice for health datasets is to identify small geographic areas and either suppress records from these small areas or aggregate them into larger ones. A recent study provided a method for deciding when an area is too small based on the uniqueness criterion. The uniqueness criterion stipulates that an the area is no longer too small when the proportion of unique individuals on the relevant variables (the quasi-identifiers) approaches zero. However, using a uniqueness value of zero is quite a stringent threshold, and is only suitable when the risks from data disclosure are quite high. Other uniqueness thresholds that have been proposed for health data are 5% and 20%.Methods: We estimated uniqueness for urban Forward Sortation Areas (FSAs) by using the 2001 long form Canadian census data representing 20% of the population. We then constructed two logistic regression models to predict when the uniqueness is greater than the 5% and 20% thresholds, and validated their predictive accuracy using 10-fold cross-validation. Predictor variables included the population size of the FSA and the maximum number of possible values on the quasi-identifiers (the number of equivalence classes).Results: All model parameters were significant and the models had very high prediction accuracy, with specificity above 0.9, and sensitivity at 0.87 and 0.74 for the 5% and 20% threshold models respectively. The application of the models was illustrated with an analysis of the Ontario newborn registry and an emergency department dataset. At the higher thresholds considerably fewer records compared to the 0% threshold would be considered to be in small areas and therefore undergo disclosure control actions. We have also included concrete guidance for data custodians in deciding which one of the three uniqueness thresholds to use (0%, 5%, 20%), depending on the mitigating controls that the data recipients have in place, the potential invasion of privacy if the data is disclosed, and the motives and capacity of the data recipient to re-identify the data.Conclusion: The models we developed can be used to manage the re-identification risk from small geographic areas. Being able to choose among three possible thresholds, a data custodian can adjust the definition of "small geographic area" to the nature of the data and recipient.