课题基金 / 基金详情

Doctoral Dissertation Research: Synthetic Data Generation for Small Area Estimation

Doctoral Dissertation Research: Synthetic Data Generation for Small Area Estimation
博士论文研究:小区域估计的综合数据生成
批准号:
0918942
负责人:
Trivellore Raghunathan
金额:
$0.6万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-15 至 2011-08-31

项目摘要

项目成果

Trivellore Raghunathan的其他基金

相似基金

相关文献

中文摘要
翻译
各种研究人员、分析人员、决策者和社区规划者对小面积估算的需求正在大幅增长,他们利用这些数据来提高对影响社区及其居民生活的问题的当前知识。统计机构定期从小地理区域收集调查数据,但往往无法以微数据形式公开发布这些数据,因为发布小区域标识符存在保密风险。本提案的主要目标是开发一种生成完全合成的微观数据集的方法,该方法允许对小区域统计数据进行有效估计,同时保护被调查者的机密性。年代数据。提出的方法将使用众所周知的贝叶斯分层建模技术,根据公共使用数据集中发现的常用变量集的假设预测模型生成模拟(或估算)值。建模方法将考虑在县和州一级发生的不同水平的变化,以便生成合成数据,为两个地理水平产生有效的推断。将考虑参数和非参数建模策略,并将基于实际数据和合成数据的小区域推断进行比较,以评估合成数据方法的效用。此外,本研究将解决在合成数据应用中通常被忽略的两个现实世界数据复杂性,包括:1)生成家庭和个人层面属性的合成数据和维护家庭内部组成结构;2)考虑复杂的样本设计特征(例如,选择、分层和聚类的不等概率)。拟议的研究将通过测试一种传播适合小区域估计的公共使用数据的替代方法,同时加强保密保护,从而开辟新的领域。如果所提议的方法证明是成功的,那么就可以避免目前要求数据用户访问受限制的数据中心设施内的小区域数据的做法。通过发布具有小区域标识符的合成微数据,用户将能够对地理级别进行定制化的小区域分析,而目前没有限制数据访问是不允许的。这一创新可能有助于满足对小区域估算的日益增长的需求,并增加利用统计机构的微数据产生的小区域估算的绝对数量。
英文摘要
Demand for small area estimates is growing heavily among a variety of researchers, analysts, decision-makers, and community planners, who use these data to advance current knowledge on issues affecting communities and the lives of their residents. Statistical agencies regularly collect survey data from small geographic areas, but are often prevented from publicly releasing these data in microdata form because of confidentiality risks associated with releasing small area identifiers. The main objective of this proposal is to develop a methodology for generating fully-synthetic micro-level datasets that permit valid estimation of small area statistics while protecting the confidentiality of respondent?s data. The proposed methodology will use well-known Bayesian hierarchical modeling techniques to generate simulated (or imputed) values based on an assumed prediction model for a commonly used set of variables found in public-use datasets. The modeling approach will account for different levels of variation occurring at the county- and state-level for the purposes of generating synthetic data that produce valid inferences for both levels of geography. Parametric and nonparametric modeling strategies will be considered, and small area inferences based on the actual data and synthetic data will be compared to evaluate the utility of the synthetic data methodology. In addition, two real-world data complexities typically ignored in synthetic data applications will be addressed in this research, including: 1) generating synthetic data for household- and individual-level attributes and maintaining the within-household composition structure; and 2) accounting for complex sample design features (e.g., unequal probabilities of selection, stratification, and clustering). The proposed research will break new ground by testing an alternative method of disseminating public-use data suitable for small area estimation while enhancing confidentiality protection. If the proposed methods prove to be successful, then the current practice of requiring data users to access small area data within restricted data center facilities may be avoided. By releasing synthetic microdata with small area identifiers, users will be able to perform customized small area analysis for levels of geography that are not currently permitted without restricted data access. This innovation may help meet the growing demand for small area estimates and increase the sheer volume of small area estimates produced using microdata from statistical agencies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Dissertation Research: An Investigation of the Nexus of Survey Nonresponse and Measurement Error
海外基金