Sharing Confidential Datasets With Geographic Identifiers Via Multiple Imputation
Sharing Confidential Datasets With Geographic Identifiers Via Multiple Imputation
批准号:
7774323
负责人:
Jerome Phillip Reiter
金额:
$19.0万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-03-01 至 2012-01-31
关键词:
AccountingAddressAgeAgingAreaComputer softwareConfidentialityConflict (Psychology)DataData LinkagesData ProtectionData QualityData SetData SourcesDatabasesDependencyEconomic ConditionsEffect Modifiers (Epidemiology)ElderlyEnvironmental Risk FactorGeographyGoalsHealthHealth Insurance Portability and Accountability ActIndividualLifeLife StyleLinkMethodologyMethodsModelingNational Research CouncilNumerical valuePoliciesPublic HealthQuality of lifeResearchResearch PersonnelRetirementRiskSimulateStatistical ModelsTailTechniquesTimebasedata sharingdensityimprovedinsightmodels and simulationpublic health relevancesocialspatial relationshipsuccesstrend
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): Geographic data can be enormously beneficial for analyses. In studies of aging, for example, they can reveal areas where elderly people live in high densities; they can illuminate how environmental factors impact the health and quality of life of elderly people; and, through contextual data, they can yield insights into the social and economic conditions and lifestyle choices of the elderly. However, geographic variables are among the most challenging data to share when making a primary data source available to others. Fine geography enables ill-intentioned users to pinpoint the identities of individuals in the shared file. Thus, data collectors typically delete or aggregate geographies to very high levels before sharing data. As examples, both deletion and aggregation are employed on geography in the public use files of the Health and Retirement Study; and, the Health Insurance Portability and Accountability Act requires that any geographic units on shared files comprise at least 20,000 people. These actions reduce the quality of analyses based on finer geographic detail, thereby sacrificing the benefits of using geography in analysis. We develop new methods to protect confidentiality in data with geographic identifiers. Our approach is to simulate values of geography and other identifying attributes, such as age, from statistical models that capture the spatial dependencies in the collected data. These simulated values replace the collected ones when sharing data. Partially simulated datasets can preserve confidentiality, since identification of units and their sensitive data is difficult when the geographies and other quasi-identifiers in the released data are not collected values. And, when the simulation models faithfully reflect the relationships in the collected data, the shared data preserve spatial associations, avoid ecological inference problems, and provide details about the tails of distributions. We have three specific aims in this proposal. First, using techniques from spatial modeling, we develop methods for simulating geographic variables conditional on attributes and for simulating at- tributes conditional on geography. Second, we apply our approach on a genuine dataset to evaluate the confidentiality protection and analytic utility of partially simulated data under three scenarios: only geography simulated, only non-geographic identifiers simulated, and both geographic and other identifiers simulated. Third, we compare our approach against aggregation techniques on the genuine dataset. Our long term goal is to develop general-purpose methodology and publicly available software for sharing inference-valid, safe data that includes finer details about geography than are currently released. This will provide statistical agencies, researchers, and other data producers with more and better options for data sharing than exist at present. PUBLIC HEALTH RELEVANCE: This research has the potential to improve the way statistical agencies, research centers, individual researchers, and other data producers share data on aging, and more broadly any health or de- mographic data containing geography. Unlike existing approaches such as deletion and high level aggregation, our approach promises to preserve fine geography and spatial relationships while pro- tecting confidentiality. Ultimately, this enables secondary data analysts to make more and better inferences, leading to deeper understanding of public health.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1080/01621459.2012.710508
发表时间:
2012-12-01
期刊:
Journal of the American Statistical Association
影响因子:
3.7
作者:
[Manrique-Vallier D, Reiter JP]
通讯作者:
Reiter JP
Multiple-Shrinkage Multinomial Probit Models with Applications to Simulating Geographies in Public Use Data.
多次收缩多项式概率模型及其在公共使用数据中模拟地理的应用。
DOI:
10.1214/13-ba816
发表时间:
2013
期刊:
Bayesian analysis
影响因子:
4.4
作者:
[Burgette,LaneF, Reiter,JeromeP]
通讯作者:
Reiter,JeromeP
DOI:
10.1002/sim.6078
发表时间:
2014-05-20
期刊:
STATISTICS IN MEDICINE
影响因子:
2
作者:
[Paiva, Thais, Chakraborty, Avishek, Reiter, Jerry, Gelfand, Alan]
通讯作者:
Gelfand, Alan
MULTIPLE IMPUTATION FOR SHARING PRECISE GEOGRAPHIES IN PUBLIC USE DATA.
用于共享公共使用数据中的精确地理信息的多重插补。
DOI:
10.1214/11-aoas506
发表时间:
2012
期刊:
The annals of applied statistics
影响因子:
--
作者:
[Wang,Hao, Reiter,JeromeP]
通讯作者:
Reiter,JeromeP
海外基金