Disclosure Limitation Methods and Information Loss for Tabular Data
Disclosure Limitation Methods and Information Loss for Tabular Data
复制标题
表格数据的披露限制方法和信息丢失
DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
S. Roehrig
中科院分区:
文献类型:
--
作者:
G. Duncan;S. Fienberg;R. Krishnan;R. Padman;S. Roehrig
Even in the age of electronic dissemination of statistical data, tables are central data products of statistical agencies. For prominent examples, see the American FactFinder (http://factfinder.census.gov/servlet/BasicFactsServlet) from the U.S. Bureau of Census, the Office of National Statistics (http://www.statistics.gov.uk/) in the U.K., and Statistics Netherlands (http://www.cbs.nl/en/figures/keyfigures/index.htm). Much survey and census data is categorical in nature and thus the representation of survey results in the form of cross-classifications or tables is a natural device for statistical reporting. But even when they collect measurement data, statistical agencies often represent the information from them in the form of discretized quantities. As a result, tables of counts represent a primary unit of reporting and analysis. Sometimes these tables represent simple cross-classifications of the counts of survey and census elements. Other times the sample units are weighted according to probabilities of selection and /or are interpretable as the numbers of people in the population (based on the sample). In such tables of counts, the occurrence of small values is usually taken to present the possibility of a disclosure risk, since data for individuals who are unique in the population may be used in matching against other databases by an intruder or data snooper.