课题基金 / 基金详情

Statistical Disclosure Limitation Methods for Tabular Data

Statistical Disclosure Limitation Methods for Tabular Data
表格数据的统计披露限制方法
批准号:
0532407
负责人:
Aleksandra Slavkovic
金额:
$26.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-10-01 至 2010-09-30

项目摘要

项目成果

Aleksandra Slavkovic的其他基金

相似基金

相关文献

中文摘要
翻译
该项目的主要目标是进一步发展适用于分析高维表格数据的跨学科理论和方法,特别注重统计披露的限制。表格数据是传播来自机密微数据的数据的主要产品,这些数据为社会科学研究提供动力,并为政策决策提供信息。随着大量表格数据的公开积累,以及记录链接方法的改进,对我们的机密性和隐私的威胁也在增加。这个项目探讨了实际问题:(1)哪些社会科学数据可以从一个小数目的表格中发布,将保持机密性?以及(2)公布的数据对统计推断有用吗?这项研究的方法论方面涉及使用边际、条件和赔率比以及对数线性模型、概率、有向无环图和代数几何中的工具来完整和不完整地描述k路列联表的概率分布。完整的规范与完全联合分发的唯一标识相关联,即充分披露,具有最大的效用。给定基于条件和边缘的任意集合的观察到的部分信息,单元格条目上的界限和分布告诉我们单元格可以采用什么值;因此可以用于风险和效用评估。这些界限可以通过线性规划和整数规划来计算。代数几何中的工具可用于计算边界和分布的归纳。该项目评估当前方法对高维离散数据的适用性和有效性;它将研究在通常是稀疏的大型表格中有条件和边际发布的披露风险和数据效用。这项研究通过研究报告比率表时的概率值四舍五入对界限清晰度的影响来改进当前的方法。该项目将开始发展一种广泛适用的理论,用于根据边际词和条件句的任意集合,在给定观察到的部分信息的情况下,评估表的空间分布。这一领域的新成果将通过引入和确认新的统计模型,推动当前统计和计算理论的前沿。直到最近,关于公布费率表对保密性的影响,人们还一无所知。发布高维列联表的条件分布对于社会科学研究人员在评估因果推理的同时仍保持机密性是有用的。这项研究的结果将进一步加强统计披露限制、离散多元统计理论和计算代数几何之间的联系。该项目提高了统计界和社会科学研究界对数据隐私问题的认识,促进了对统计披露限制的研究,并招募年轻学者研究适用于社会科学和行为科学的新统计方法。这项研究为政府机构和公共卫生研究人员提供了评估高维表格数据发布的安全性和实用性的新工具。这些工具和数据发布补充了目前公共使用数据文件的传播做法,并在分享数据和研究成果方面提供了更大的灵活性,同时保持机密性。该奖项是作为2005财政年度数学科学优先领域数学社会科学和行为科学特别竞赛的一部分得到支持的。
英文摘要
The main goal of this project is to further develop interdisciplinary theory and methodology applicable for analysis of high-dimensional tabular data, with a special focus on statistical disclosure limitation. Tabular data are a staple product in disseminating data derived from confidential microdata that fuels social science research and informs policy decisions. As the amount of tabular data accumulates in public, and record linkage methodologies improve, so does the threat to our confidentiality and privacy. This project explores the practical questions: (1) What social science data, releasable from a table with small counts, will maintain confidentiality? and (2) Will the released data be useful for statistical inference? The methodological aspects of this research deal with complete and incomplete characterizations of probability distributions for k-way contingency tables using marginals, conditionals, and odds ratios, and tools from log-linear models, probability, directed acyclic graphs, and algebraic geometry. Complete specifications are associated with unique identification of the full joint distribution, i.e., full disclosure, with the maximum utility. Given observed partial information based on an arbitrary collection of conditionals and marginals, the bounds and distributions on the cell entries tell us what values the cells can take; thus can be used for risk and utility assessment. The bounds can be calculated via linear and integer programming. Tools from algebraic geometry can be used for the calculation of bounds and induction of distributions. This project evaluates the applicability and effectiveness of the current methodology for high-dimensional discrete data; it will study the disclosure risk and data utility of conditional and marginal releases in large often sparse tables. This research improves the current methodology by studying the effects of rounding of probability values when reporting tables of rates on sharpness of bounds. The project will start developing a broadly applicable theory for assessing distributions over the space of tables given observed partial information based on arbitrary collections of marginals and conditionals. New results in this area will advance the current frontier of statistical and computational theory by introducing and confirming new statistical models. Until recently nothing was known about the effects on confidentiality of releasing tables of rates. Releasing conditional distributions for high-dimensional contingency tables could be useful for social science researchers assessing causal inference while still maintaining confidentiality. The results of this research will further the connections between statistical disclosure limitation, discrete multivariate statistical theory, and computational algebraic geometry. The project increases awareness of data privacy issues in both the statistics and social sciences research communities, promotes research on statistical disclosure limitation, and recruits young scholars to study new statistical methods applicable to social and behavioral sciences. This research provides government agencies and public health researchers with new tools for evaluating the safety and utility of high-dimensional tabular data releases. These tools and data releases complement current dissemination practices of public use data files, and offer more flexibility in sharing data and research results while maintaining confidentiality. This award was supported as part of the fiscal year 2005 Mathematical Sciences priority area special competition on Mathematical Social and Behavioral Sciences (MSBS).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Formal Privacy for Complex Data Objects
Collaborative Research: Record Linkage and Privacy-Preserving Methods for Big Data
CDI-Type II: Collaborative Research: Integrating Statistical and Computational Approaches to Privacy
海外基金