课题基金 / 基金详情

Sharing Confidential Datasets With Geographic Identifiers Via Multiple Imputation

Sharing Confidential Datasets With Geographic Identifiers Via Multiple Imputation
通过多重插补与地理标识符共享机密数据集
批准号:
7774323
负责人:
Jerome Phillip Reiter
金额:
$19.0万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-03-01 至 2012-01-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
描述(申请人提供):地理数据对分析有极大的帮助。例如,在老龄化研究中,它们可以揭示老年人高密度生活的地区;它们可以阐明环境因素如何影响老年人的健康和生活质量;通过背景数据,它们可以洞察老年人的社会和经济状况以及生活方式选择。然而,在向其他人提供主数据源时,地理变量是共享最具挑战性的数据之一。精细的地理位置使恶意用户能够确定共享文件中个人的身份。因此,数据收集者通常会在共享数据之前删除地理位置或将地理位置聚合到非常高的级别。例如,在健康和退休研究的公共使用文件中,删除和聚合都是按地理位置进行的;《健康保险可携带性和责任法》要求共享文件中的任何地理单位至少包括20,000人。这些操作降低了基于更精细的地理细节的分析质量,从而牺牲了在分析中使用地理的好处。我们开发了新的方法来保护具有地理标识的数据的机密性。我们的方法是从统计模型中模拟地理和其他识别属性的值,例如年龄,这些统计模型捕捉收集的数据中的空间相关性。在共享数据时,这些模拟值将替换收集的值。部分模拟的数据集可以保持机密性,因为当发布的数据中的地理位置和其他准标识符不是采集值时,识别单元及其敏感数据是困难的。并且,当模拟模型真实地反映了收集的数据中的关系时,共享数据保留了空间关联,避免了生态推断问题,并提供了关于分布尾部的细节。我们在这项提案中有三个具体目标。首先,利用空间建模技术,我们发展了以属性为条件的地理变量模拟方法和以地理条件为条件的属性模拟方法。其次,我们在真实数据集上对部分模拟数据在三种场景下的机密性保护和分析效用进行了评估:仅模拟地理标识、仅模拟非地理标识以及同时模拟地理标识和其他标识。第三,我们将我们的方法与真实数据集上的聚合技术进行了比较。我们的长期目标是开发通用的方法和公开可用的软件来共享推理--有效的、安全的数据,其中包括比目前发布的更详细的地理信息。这将为统计机构、研究人员和其他数据生产者提供比目前更多和更好的数据共享选择。与公共健康相关:这项研究有可能改善统计机构、研究中心、个人研究人员和其他数据生产者共享老龄化数据的方式,更广泛地说,共享任何包含地理位置的健康或运动数据。与现有的删除和高级聚合等方法不同,我们的方法承诺在保护机密性的同时保留良好的地理和空间关系。最终,这使二级数据分析人员能够做出更多、更好的推断,从而加深对公共卫生的理解。
英文摘要
DESCRIPTION (provided by applicant): Geographic data can be enormously beneficial for analyses. In studies of aging, for example, they can reveal areas where elderly people live in high densities; they can illuminate how environmental factors impact the health and quality of life of elderly people; and, through contextual data, they can yield insights into the social and economic conditions and lifestyle choices of the elderly. However, geographic variables are among the most challenging data to share when making a primary data source available to others. Fine geography enables ill-intentioned users to pinpoint the identities of individuals in the shared file. Thus, data collectors typically delete or aggregate geographies to very high levels before sharing data. As examples, both deletion and aggregation are employed on geography in the public use files of the Health and Retirement Study; and, the Health Insurance Portability and Accountability Act requires that any geographic units on shared files comprise at least 20,000 people. These actions reduce the quality of analyses based on finer geographic detail, thereby sacrificing the benefits of using geography in analysis. We develop new methods to protect confidentiality in data with geographic identifiers. Our approach is to simulate values of geography and other identifying attributes, such as age, from statistical models that capture the spatial dependencies in the collected data. These simulated values replace the collected ones when sharing data. Partially simulated datasets can preserve confidentiality, since identification of units and their sensitive data is difficult when the geographies and other quasi-identifiers in the released data are not collected values. And, when the simulation models faithfully reflect the relationships in the collected data, the shared data preserve spatial associations, avoid ecological inference problems, and provide details about the tails of distributions. We have three specific aims in this proposal. First, using techniques from spatial modeling, we develop methods for simulating geographic variables conditional on attributes and for simulating at- tributes conditional on geography. Second, we apply our approach on a genuine dataset to evaluate the confidentiality protection and analytic utility of partially simulated data under three scenarios: only geography simulated, only non-geographic identifiers simulated, and both geographic and other identifiers simulated. Third, we compare our approach against aggregation techniques on the genuine dataset. Our long term goal is to develop general-purpose methodology and publicly available software for sharing inference-valid, safe data that includes finer details about geography than are currently released. This will provide statistical agencies, researchers, and other data producers with more and better options for data sharing than exist at present. PUBLIC HEALTH RELEVANCE: This research has the potential to improve the way statistical agencies, research centers, individual researchers, and other data producers share data on aging, and more broadly any health or de- mographic data containing geography. Unlike existing approaches such as deletion and high level aggregation, our approach promises to preserve fine geography and spatial relationships while pro- tecting confidentiality. Ultimately, this enables secondary data analysts to make more and better inferences, leading to deeper understanding of public health.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1080/01621459.2012.710508
发表时间: 2012-12-01
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Manrique-Vallier D, Reiter JP]
通讯作者: Reiter JP
Multiple-Shrinkage Multinomial Probit Models with Applications to Simulating Geographies in Public Use Data.
多次收缩多项式概率模型及其在公共使用数据中模拟地理的应用。
DOI: 10.1214/13-ba816
发表时间: 2013
期刊: Bayesian analysis
影响因子: 4.4
作者: [Burgette,LaneF, Reiter,JeromeP]
通讯作者: Reiter,JeromeP
DOI: 10.1002/sim.6078
发表时间: 2014-05-20
期刊: STATISTICS IN MEDICINE
影响因子: 2
作者: [Paiva, Thais, Chakraborty, Avishek, Reiter, Jerry, Gelfand, Alan]
通讯作者: Gelfand, Alan
MULTIPLE IMPUTATION FOR SHARING PRECISE GEOGRAPHIES IN PUBLIC USE DATA.
用于共享公共使用数据中的精确地理信息的多重插补。
DOI: 10.1214/11-aoas506
发表时间: 2012
期刊: The annals of applied statistics
影响因子: --
作者: [Wang,Hao, Reiter,JeromeP]
通讯作者: Reiter,JeromeP
海外基金