课题基金 / 基金详情

Methodology for Improving Public Use Data Dissemination Via Multiply-Imputed, Partially Synthetic Data

Methodology for Improving Public Use Data Dissemination Via Multiply-Imputed, Partially Synthetic Data
通过多重插补、部分合成数据改进公共使用数据传播的方法
批准号:
0751671
负责人:
Jerome Reiter
金额:
$18.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-06-01 至 2011-05-31

项目摘要

项目成果

Jerome Reiter的其他基金

相似基金

相关文献

中文摘要
翻译
向公众传播数据的统计机构和其他组织在道德上和法律上往往需要保护答卷人身份和敏感属性的机密性。 为了满足这些要求,各机构可以发布多重估算的、部分合成的数据。 这些单位包括最初调查时使用某些数值的单位,例如被多重估算所取代的高披露风险的敏感数值或关键标识的数值。本研究通过解决实施过程中的四个关键问题,提高了部分合成数据方法的风险效用。 首先,本研究开发了量化部分合成数据集识别披露风险的方法。 这些措施占(i)存在于所有的合成数据集的信息,(ii)各种假设入侵者的知识和行为,以及(iii)发布的合成数据生成模型的细节。 这些信息对于试图评估合成数据所提供的保护的数据生产者至关重要。 其次,研究提供了数据生产者可以用来选择要合成的值的策略。 这些策略优化了候选值集的风险和效用之间的权衡。第三,该研究产生了选择合成数据集的策略。 例如,数据生产者可以丢弃披露风险过高或数据效用过低的合成数据集。 该研究产生的指导方针,这种选择如何影响使用现有的方法作出的推断,并制定适当的推理方法的情况下,选择的影响是巨大的。 最后,研究开发灵活的,非参数化的建模策略,用于基于机器学习技术的合成数据生成。 本研究为联邦机构、调查组织、研究中心和其他数据生产者提供了比目前更多更好的公共使用数据传播选择。 随着恶意数据用户可用资源的不断扩展,使用传统的披露限制技术保护公共使用数据所需的更改-例如交换数据值,添加随机噪声或聚合数据-可能变得如此极端,以至于对于许多分析来说,发布的数据不再有用。 另一方面,合成数据有可能使公众能够使用数据传播,同时保持数据效用。 最终,有了更高质量的公共使用数据,二级数据分析师可以做出更多更好的推断,从而加深对社会科学和政策问题的理解。
英文摘要
Statistical agencies and other organizations that disseminate data to the public are ethically and often legally required to protect the confidentiality of respondents' identities and sensitive attributes. To satisfy these requirements, agencies can release multiply-imputed, partially synthetic data. These comprise the units originally surveyed with some values, such as sensitive values at high risk of disclosure or values of key identifiers, replaced with multiple imputations. This research improves the risk-utility profile of partially synthetic data approaches by addressing four key issues in their implementation. First, the research develops methods for quantifying identification disclosure risks for partially synthetic data sets. These measures account for (i) the information existing in all the synthetic data sets, (ii) various assumptions about intruder knowledge and behavior, and (iii) the details released about the synthetic data generation model. This information is crucial to data producers seeking to evaluate the protection afforded by synthetic data. Second, the research provides strategies that data producers can use to select values to synthesize. The strategies optimize the trade-offs between risk and utility for candidate sets of values. Third, the research yields strategies for selecting synthetic data sets. For example, the data producer can throw out synthetic data sets that are too high in disclosure risk or too low in data utility. The research produces guidelines for how such selection impacts inferences made using existing methods, and it develops appropriate methods of inference for situations where the effects of selection are substantial. Finally, the research develops flexible, nonparametric modeling strategies for synthetic data generation based on techniques from machine learning. This improves the analytic validity of partially synthetic data approaches.This research provides federal agencies, survey organizations, research centers, and other data producers with more and better options for public use data dissemination than exist at present. As resources available to malicious data users continue to expand, the alterations needed to protect public use data with traditional disclosure limitation techniques---such as swapping data values, adding random noise, or aggregating data---may become so extreme that, for many analyses, the released data are no longer useful. Synthetic data, on the other hand, have the potential to enable public use data dissemination while preserving data utility. Ultimately, with higher quality public use data, secondary data analysts can make more and better inferences, leading to deeper understanding of social science and policy questions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enhancing Synthetic Data Techniques for Practical Applications
  • 批准号:
    2217456
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2022
  • 负责人:
    Jerome Reiter
  • 依托单位:
Leveraging Auxiliary Information on Marginal Distributions in Multiple Imputation for Survey Nonresponse
  • 批准号:
    1733835
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2017
  • 负责人:
    Jerome Reiter
  • 依托单位:
CIF21 DIBBs: An Integrated System for Public/Private Access to Large-Scale, Confidential Social Science Data
  • 批准号:
    1443014
  • 项目类别:
    Standard Grant
  • 资助金额:
    $149.87万
  • 财政年份:
    2015
  • 负责人:
    Jerome Reiter
  • 依托单位:
NCRN-MN: Triangle Census Research Network
  • 批准号:
    1131897
  • 项目类别:
    Standard Grant
  • 资助金额:
    $299.76万
  • 财政年份:
    2011
  • 负责人:
    Jerome Reiter
  • 依托单位:
国内基金
海外基金
Improving modelling of compact binary evolution.
  • 批准号:
    10903001
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2009
  • 负责人:
    史蒂芬
  • 依托单位: