Methodology for Improving Public Use Data Dissemination Via Multiply-Imputed, Partially Synthetic Data
Methodology for Improving Public Use Data Dissemination Via Multiply-Imputed, Partially Synthetic Data
批准号:
0751671
负责人:
Jerome Reiter
金额:
$18.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-06-01 至 2011-05-31
中文摘要
向公众传播数据的统计机构和其他组织在道德上和法律上往往被要求保护答复者身份和敏感属性的机密性。为了满足这些要求,机构可以发布多个推算的、部分合成的数据。这些单位包括最初调查的一些值,例如具有高披露风险的敏感值或关键识别符的值,由多个推算取代。这项研究通过解决实施中的四个关键问题,改善了部分合成数据方法的风险-效用概况。首先,该研究开发了量化部分合成数据集的身份披露风险的方法。这些措施考虑了(I)所有合成数据集中存在的信息,(Ii)关于入侵者知识和行为的各种假设,以及(Iii)关于合成数据生成模型的发布的细节。这些信息对于寻求评估合成数据所提供的保护的数据生产者来说至关重要。其次,这项研究提供了数据生产者可以用来选择要综合的值的策略。这些策略优化了候选值集合的风险和效用之间的权衡。第三,研究得出了选择合成数据集的策略。例如,数据生产者可以丢弃披露风险太高或数据利用率太低的合成数据集。这项研究为这种选择如何影响使用现有方法进行的推理制定了指导方针,并为选择的影响很大的情况开发了适当的推理方法。最后,研究开发了灵活的、基于机器学习技术的合成数据生成的非参数建模策略。这提高了部分合成数据方法的分析有效性。这项研究为联邦机构、调查组织、研究中心和其他数据生产者提供了比目前更多和更好的公共使用数据传播选择。随着恶意数据用户可用的资源不断扩大,使用传统披露限制技术保护公共使用数据所需的更改-例如交换数据值、添加随机噪声或聚合数据--可能会变得如此极端,以至于对许多分析来说,已发布的数据不再有用。另一方面,合成数据有可能在保留数据效用的同时,使公众能够传播数据。最终,有了更高质量的公共使用数据,二次数据分析师可以做出更多更好的推断,从而加深对社会科学和政策问题的理解。
英文摘要
Statistical agencies and other organizations that disseminate data to the public are ethically and often legally required to protect the confidentiality of respondents' identities and sensitive attributes. To satisfy these requirements, agencies can release multiply-imputed, partially synthetic data. These comprise the units originally surveyed with some values, such as sensitive values at high risk of disclosure or values of key identifiers, replaced with multiple imputations. This research improves the risk-utility profile of partially synthetic data approaches by addressing four key issues in their implementation. First, the research develops methods for quantifying identification disclosure risks for partially synthetic data sets. These measures account for (i) the information existing in all the synthetic data sets, (ii) various assumptions about intruder knowledge and behavior, and (iii) the details released about the synthetic data generation model. This information is crucial to data producers seeking to evaluate the protection afforded by synthetic data. Second, the research provides strategies that data producers can use to select values to synthesize. The strategies optimize the trade-offs between risk and utility for candidate sets of values. Third, the research yields strategies for selecting synthetic data sets. For example, the data producer can throw out synthetic data sets that are too high in disclosure risk or too low in data utility. The research produces guidelines for how such selection impacts inferences made using existing methods, and it develops appropriate methods of inference for situations where the effects of selection are substantial. Finally, the research develops flexible, nonparametric modeling strategies for synthetic data generation based on techniques from machine learning. This improves the analytic validity of partially synthetic data approaches.This research provides federal agencies, survey organizations, research centers, and other data producers with more and better options for public use data dissemination than exist at present. As resources available to malicious data users continue to expand, the alterations needed to protect public use data with traditional disclosure limitation techniques---such as swapping data values, adding random noise, or aggregating data---may become so extreme that, for many analyses, the released data are no longer useful. Synthetic data, on the other hand, have the potential to enable public use data dissemination while preserving data utility. Ultimately, with higher quality public use data, secondary data analysts can make more and better inferences, leading to deeper understanding of social science and policy questions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enhancing Synthetic Data Techniques for Practical Applications
-
批准号:2217456
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2022
-
负责人:Jerome Reiter
-
依托单位:
Leveraging Auxiliary Information on Marginal Distributions in Multiple Imputation for Survey Nonresponse
-
批准号:1733835
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2017
-
负责人:Jerome Reiter
-
依托单位:
CIF21 DIBBs: An Integrated System for Public/Private Access to Large-Scale, Confidential Social Science Data
-
批准号:1443014
-
项目类别:Standard Grant
-
资助金额:$149.87万
-
财政年份:2015
-
负责人:Jerome Reiter
-
依托单位:
NCRN-MN: Triangle Census Research Network
-
批准号:1131897
-
项目类别:Standard Grant
-
资助金额:$299.76万
-
财政年份:2011
-
负责人:Jerome Reiter
-
依托单位:
Multiple Imputation Methods for Handling Missing Data in Longitudinal Studies with Refreshment Samples
-
批准号:1061241
-
项目类别:Standard Grant
-
资助金额:$16.0万
-
财政年份:2011
-
负责人:Jerome Reiter
-
依托单位:
TC: Large: Collaborative Research: Practical Privacy: Metrics and Methods for Protecting Record-level and Relational Data
-
批准号:1012141
-
项目类别:Continuing Grant
-
资助金额:$58.32万
-
财政年份:2010
-
负责人:Jerome Reiter
-
依托单位:
国内基金
海外基金
Improving modelling of compact binary evolution.
-
批准号:10903001
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2009
-
负责人:史蒂芬
-
依托单位: