Disclosure Limitation Methods and Information Loss for Tabular Data

Disclosure Limitation Methods and Information Loss for Tabular Data
复制标题

表格数据的披露限制方法和信息丢失

DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
S. Roehrig
S. Roehrig
中科院分区:
--
文献类型:
--
作者:
G. Duncan;S. Fienberg;R. Krishnan;R. Padman;S. Roehrig

文献摘要

被引文献

相似文献

即使在统计数据电子传播的时代,表格也是统计机构的核心数据产品。有关突出示例,请参阅美国人口普查局的 American FactFinder (http://factfinder.census.gov/servlet/BasicFactsServlet)、英国国家统计局 (http://www.statistics.gov.uk/) 和荷兰统计局 (http://www.cbs.nl/en/figures/keyfigures/index.htm)。许多调查和普查数据本质上是分类的,因此以交叉分类或表格的形式表示调查结果是统计报告的自然手段。但即使在收集测量数据时,统计机构也常常以离散量的形式表示来自测量数据的信息。因此,计数表代表了报告和分析的主要单位。有时,这些表格代表调查和普查要素计数的简单交叉分类。其他时候,样本单位根据选择概率进行加权和/或可解释为总体中的人数(基于样本)。在此类计数表中,小值的出现通常被认为存在泄露风险的可能性,因为入侵者或数据窥探者可能会使用群体中唯一的个体的数据来与其他数据库进行匹配。
Even in the age of electronic dissemination of statistical data, tables are central data products of statistical agencies. For prominent examples, see the American FactFinder (http://factfinder.census.gov/servlet/BasicFactsServlet) from the U.S. Bureau of Census, the Office of National Statistics (http://www.statistics.gov.uk/) in the U.K., and Statistics Netherlands (http://www.cbs.nl/en/figures/keyfigures/index.htm). Much survey and census data is categorical in nature and thus the representation of survey results in the form of cross-classifications or tables is a natural device for statistical reporting. But even when they collect measurement data, statistical agencies often represent the information from them in the form of discretized quantities. As a result, tables of counts represent a primary unit of reporting and analysis. Sometimes these tables represent simple cross-classifications of the counts of survey and census elements. Other times the sample units are weighted according to probabilities of selection and /or are interpretable as the numbers of people in the population (based on the sample). In such tables of counts, the occurrence of small values is usually taken to present the possibility of a disclosure risk, since data for individuals who are unique in the population may be used in matching against other databases by an intruder or data snooper.