General and specific utility measures for synthetic data

General and specific utility measures for synthetic data
复制标题

DOI:
10.1111/rssa.12358
复制
发表时间:
2018-06-01
影响因子:
2
通讯作者:
Slavkovic, Aleksandra
Slavkovic, Aleksandra
中科院分区:
数学4区
文献类型:
--
作者:
Snoke, Joshua;Raab, Gillian M.;Slavkovic, Aleksandra

文献摘要

被引文献

相似文献

当对可能的披露的担忧限制了原始记录的可用性时,数据持有者可以制作数据集的合成版本。本文关注的是判断这种合成数据是否具有与原始数据相比较的分布的方法:我们称之为通用工具。我们考虑一般效用与特定效用的比较:从合成数据和原始数据分析结果的相似性。我们适应以前的一般措施的数据效用,倾向得分均方误差pMSE,合成数据的特定情况下,并推导出其分布的情况下,正确的合成模型被用来创建合成数据。我们的渐近结果证实了模拟研究。我们还考虑了两个具体的效用措施,置信区间重叠和标准化差异的汇总统计量,我们与一般的效用结果进行比较。我们提出了两个对比的数据合成的例子:一个是说明合成数据,被评估为有用的一般和具体的措施,第二个都不是这种情况。对于第二种情况下,我们展示了一般效用措施如何识别合成数据的不足之处,并建议如何告知可能的改进合成方法。
Data holders can produce synthetic versions of data sets when concerns about potential disclosure restrict the availability of the original records. The paper is concerned with methods to judge whether such synthetic data have a distribution that is comparable with that of the original data: what we term general utility. We consider how general utility compares with specific utility: the similarity of results of analyses from the synthetic data and the original data. We adapt a previous general measure of data utility, the propensity score mean-squared error pMSE, to the specific case of synthetic data and derive its distribution for the case when the correct synthesis model is used to create the synthetic data. Our asymptotic results are confirmed by a simulation study. We also consider two specific utility measures, confidence interval overlap and standardized difference in summary statistics, which we compare with the general utility results. We present two contrasting examples of data syntheses: one illustrating synthetic data that is evaluated as being useful by both general and specific measures and the second where neither is the case. For the second case we show how the general utility measures can identify the deficiencies of the synthetic data and suggest how this can inform possible improvements to the synthesis method.