DataSifter II: Partially synthetic data sharing of sensitive information containing time-varying correlated observations.

DataSifter II: Partially synthetic data sharing of sensitive information containing time-varying correlated observations.
复制标题

DOI:
10.1177/17483026211065379
复制
发表时间:
2022
影响因子:
0.9
通讯作者:
Dinov, Ivo D.
Dinov, Ivo D.
中科院分区:
其他
文献类型:
--
作者:
Zhou, Nina;Wang, Lu;Marino, Simeone;Zhao, Yi;Dinov, Ivo D.

文献摘要

参考文献

相似文献

公众对利用汇总的敏感信息进行快速数据驱动的科学调查的需求很大。然而,许多技术挑战和监管政策阻碍了有效的数据共享。在这项研究中,我们描述了一种部分合成的数据生成技术,用于创建匿名数据档案,其联合分布非常类似于原始(敏感)数据。具体来说,我们介绍了时变相关数据的DataSifter技术(DataSifter II),它依赖于使用广义线性混合模型和随机效应期望最大化树的基于迭代模型的插补。DataSifter II可用于生成用于测试和验证新分析技术的合成重复测量数据。与多重插补方法相比,DataSifter II在模拟和真实的临床数据上的应用表明,新方法可大幅降低重新识别风险(数据隐私),同时保留混淆数据中的分析价值(数据实用性)。DataSifter II在模拟中的表现涉及数据中20%的人为缺失,与多重插补方法相比,披露风险至少降低了80%,而对数据分析值没有实质性影响。在单独的临床数据(重症监护III的医疗信息市场)验证中,从原始数据中得出的基于模型的统计推断与使用DataSifter II模糊(筛选)数据获得的类似分析推断一致。对于包含敏感信息的大型时变数据集,所提出的技术提供了一种自动化工具,用于减轻数据共享的障碍,并促进有效,先进和协作的分析。
There is a significant public demand for rapid data-driven scientific investigations using aggregated sensitive information. However, many technical challenges and regulatory policies hinder efficient data sharing. In this study, we describe a partially synthetic data generation technique for creating anonymized data archives whose joint distributions closely resemble those of the original (sensitive) data. Specifically, we introduce the DataSifter technique for time-varying correlated data (DataSifter II), which relies on an iterative model-based imputation using generalized linear mixed model and random effects-expectation maximization tree. DataSifter II can be used to generate synthetic repeated measures data for testing and validating new analytical techniques. Compared to the multiple imputation method, DataSifter II application on simulated and real clinical data demonstrates that the new method provides extensive reduction of re-identification risk (data privacy) while preserving the analytical value (data utility) in the obfuscated data. The performance of the DataSifter II on a simulation involving 20% artificially missingness in the data, shows at least 80% reduction of the disclosure risk, compared to the multiple imputation method, without a substantial impact on the data analytical value. In a separate clinical data (Medical Information Mart for Intensive Care III) validation, a model-based statistical inference drawn from the original data agrees with an analogous analytical inference obtained using the DataSifter II obfuscated (sifted) data. For large time-varying datasets containing sensitive information, the proposed technique provides an automated tool for alleviating the barriers of data sharing and facilitating effective, advanced, and collaborative analytics.
DOI: 10.1186/gm316
发表时间: 2012-02-27
期刊: Genome medicine
影响因子: 12.3
作者:
Caulfield T;Harmon SH;Joly Y
通讯作者: Joly Y
DOI: 10.7554/elife.16800
发表时间: 2016-07-07
期刊: ELIFE
影响因子: 7.7
作者:
McKiernan, Erin C.;Bourne, Philip E.;Yarkoni, Tal
通讯作者: Yarkoni, Tal
DOI: 10.1007/s10586-018-2723-9
发表时间: 2019-07-01
影响因子: 4.4
作者:
Kanna, G. Prabu;Vasudevan, V.
通讯作者: Vasudevan, V.
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.2105/ajph.2014.302406
发表时间: 2015-05-01
影响因子: 12.7
作者:
Keegan, Theresa H. M.;Kurian, Allison W.;Gomez, Scarlett L.
通讯作者: Gomez, Scarlett L.