课题基金 / 基金详情

Confidentiality and Estimation for Large Sparse Multi-Dimensional Contingency Tables

Confidentiality and Estimation for Large Sparse Multi-Dimensional Contingency Tables
大型稀疏多维列联表的保密性和估计
批准号:
0631589
负责人:
Stephen Fienberg
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-09-15 至 2011-08-31

项目摘要

项目成果

Stephen Fienberg的其他基金

相似基金

相关文献

中文摘要
翻译
这项研究项目涉及使用大型稀疏列联表的两个关键方面:在与其他研究人员共享数据时保护响应的机密性,以及稀疏性对对数线性模型中的最大似然估计的影响。第一个问题需要评估与从分类数据库部分发布信息有关的披露风险,例如,以涉及变量子集的边际表格的形式。第二个问题涉及开发适用于稀疏分类数据的对数线性模型分析中的模型选择和估计/检验的通用推理方法。这些看似独立的问题之间的联系源于代数统计的常见统计和数学形式主义。这项研究将产生新的计算算法和可共享的计算机代码,供行为和社会科学研究人员使用,以及将使用最大似然和对数线性模型的单元估计问题与保密保护联系起来的基本方法和理论。这项活动的预期成果将包括:(1)对行为和社会科学数据进行量化分析和解释以及确定披露风险的更有效的推论程序;(2)以广大从业人员和研究人员为对象分析分类数据的统计软件,该软件将以计算机源代码和模块化可执行文件的形式开发和免费分发;(3)评估与发布边际总数有关的披露风险的更有效的数字程序。对数-线性模型分析形成了一套成熟和强大的统计工具,用于研究分类数据,特别是以多维交叉分类或多向联想表的形式,事实证明,这些模型对于分析来自社会科学和行为科学的许多领域以及其他科学领域的数据是必不可少的。例如,在典型的抽样调查中,为数千人生成了关于大量分类变量的数据,以衡量就业、收入、健康状况等方面的信息。由此产生的这些变量的交叉分类很大,即涉及数千个单元格,而且很稀疏,即大多数单元格条目要么非常小,要么包含零计数。类似的问题也出现在社交网络的研究、公共卫生和医学以及遗传学数据库的分析中。代数几何数学领域的最新发展为表示与这种列联表数据相关的对数线性模型提供了一种新颖而强大的形式。本项目将利用这一数学形式侧重于大型稀疏列联表的两个不同方面:(1)在与其他用户共享数据时保护数据提供者的隐私,同时(2)通过开发新的对数线性模型计算方法,确保这些表可用于统计分析。该项目的成果将改善获取数据进行二次分析的机会,并提高研究人员和分析人员利用大型稀疏数据库中信息的能力。
英文摘要
This research project deals with two crucial aspects of working with large sparse contingency tables: protecting the confidentiality of responses when data are shared with other researchers, and the implications of sparsity for maximum likelihood estimation in log-linear models. The first problem entails the evaluation of the disclosure risk associated with the partial release of information from a classified database, e.g., in the form of marginal tables involving subsets of variables. The second problem is concerned with developing general-purpose inferential methodologies for model selection and estimation/testing in log-linear model analysis that are appropriate for sparse categorical data. The links between these seemingly separate problems emanate from the common statistical and mathematical formalism of algebraic statistics. This research will produce new computational algorithms and sharable computer code for use by behavioral and social science researchers, as well as foundational methods and theory linking the problems of cell estimation using maximum likelihood and log-linear models and confidentiality protection. The expected outcomes of this activity will include: (1) more effective inferential procedures for the quantitative analysis and interpretation of behavioral and social science data and for the determination of the risk of disclosure; (2) statistical software for the analysis of categorical data targeted at a large audience of practitioners and researchers, which will be developed and freely distributed in the form of both computer source codes and modular, executable files; (3) more efficient numerical procedures for assessing the disclosure risk associated with the release of marginal totals.Log-linear models analysis forms a well-established and powerful set of statistical tools for the study of categorical data, especially in the form of multi-dimentional cross-classifications or multi-way contingency tables, These models have proved to be essential for the analysis of data emanating from many areas of the social and behavioral sciences, as well as in other scientific areas. For example, in a typical sample survey,data are generated for several thousand individuals on a large number of categorical variables, measuring such information on employment, income, health status, etc. The resulting cross-classification of these variables is large, i.e., involving many thousands of cells, and sparse, i.e., most of the cell entries are either very small or contain zero counts. Similar problems arise in the study of social networks, in public health and medicine, and in the analysis of genetics databases. Recent developments in the mathematical area of algebraic geometry have provided a novel and powerful formalism for the representation of log-linear models relevant for such contingency table data. This project will use this mathematical formalism to focus on two different aspects of large sparse contingency tables: (1) Protecting the privacy of the data providers when data are shared with other users, while at the same time (2) Ensuring that such tables are useful for statistical analysis by developing new methods for log-linear model computation. The results of the project will improve access to data for secondary analysis and enhance the capacity of researchers and analysts to exploit the information in large sparse databases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CDI-Type II: Collaborative Research: Integrating Statistical and Computational Approaches to Privacy
  • 批准号:
    0941518
  • 项目类别:
    Standard Grant
  • 资助金额:
    $61.5万
  • 财政年份:
    2010
  • 负责人:
    Stephen Fienberg
  • 依托单位:
Participant Support for Workshop on Statistical Methods for the Analysis of Network Data in in Dublin, Ireland.
  • 批准号:
    0924358
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2009
  • 负责人:
    Stephen Fienberg
  • 依托单位:
Travel Grant Proposal for Workshop on Data Confidentiality
  • 批准号:
    0741571
  • 项目类别:
    Standard Grant
  • 资助金额:
    $3.1万
  • 财政年份:
    2007
  • 负责人:
    Stephen Fienberg
  • 依托单位:
Workshop on Privacy and Confidentiality; July 19-26, 2005,Italy.
  • 批准号:
    0517956
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.5万
  • 财政年份:
    2005
  • 负责人:
    Stephen Fienberg
  • 依托单位:
海外基金