Practical Data Synthesis for Large Samples

Practical Data Synthesis for Large Samples
复制标题

大样本的实用数据合成

DOI:
10.29012/jpc.v7i3.407
复制
发表时间:
2018
期刊:
J. Priv. Confidentiality
影响因子:
--
通讯作者:
C. Dibben
C. Dibben
中科院分区:
--
文献类型:
--
作者:
G. Raab;B. Nowok;C. Dibben

文献摘要

被引文献

相似文献

我们描述了创建和使用合成数据的结果,这些数据是在一个项目的背景下得出的,该项目旨在为英国纵向研究的用户提供合成提取物。对现有的大型综合数据集推理方法进行了评述。我们引入了新的方差估计,用于完全合成数据的大样本,这些数据不需要从观察数据得出的后验预测分布中生成,并且可以与单个合成数据集一起使用。我们就如何根据这些结果综合数据提出建议。这些结果的实际后果用苏格兰纵向研究的一个例子来说明。
We describe results on the creation and use of synthetic data that were derived in the context of a project to make synthetic extracts available for users of the UK Longitudinal Studies. A critical review of existing methods of inference from large synthetic data sets is presented. We introduce new variance estimates for use with large samples of completely synthesised data that do not require them to be generated from the posterior predictive distribution derived from the observed data and can be used with a single synthetic data set. We make recommendations on how to synthesise data based on these results. The practical consequences of these results are illustrated with an example from the Scottish Longitudinal Study.