Using Statistics to Protect Privacy

Using Statistics to Protect Privacy
复制标题

使用统计数据保护隐私

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Jerome P. Reiter
Jerome P. Reiter
中科院分区:
--
文献类型:
--
作者:
A. Karr;Jerome P. Reiter

文献摘要

参考文献

被引文献

相似文献

1.产生数据的机构-例如官方统计机构、调查组织和主要调查人员,以下统称为机构-长期以来一直向研究人员、政策分析人员、决策者和公众提供其数据。与此同时,这些机构在道德上和法律上都有义务保护数据主体身份和敏感属性的机密性。简单地剥离名称、确切地址和其他直接标识符通常不足以保护机密性。当发布的数据包括外部文件中容易获得的变量时,例如人口统计特征或就业历史,恶意用户-此后称为入侵者-可能能够将发布数据中的记录与外部文件中的记录联系起来,从而损害了机构对提供数据者的保密承诺。为了应对这一威胁,各机构已经制定了各种各样的战略,以减少意外披露的风险,从限制数据访问到在发布前更改数据。属于后一类的策略被称为统计披露限制(SDL)技术。大多数SDL技术都是针对概率调查或人口普查中的数据开发的。即使是完整的形式,这些数据通常也不会被认为是大数据,就规模(案例和属性的数量),属性类型的复杂性或结构而言:大多数数据集都是作为平面文件发布的,如果不是实际结构化的话。
Introduction Those who generate data – for example, official statistics agencies, survey organizations, and principal investigators, henceforth all called agencies – have a long history of providing access to their data to researchers, policy analysts, decision makers, and the general public. At the same time, these agencies are obligated ethically and often legally to protect the confidentiality of data subjects’ identities and sensitive attributes. Simply stripping names, exact addresses, and other direct identifiers typically does not suffice to protect confidentiality. When the released data include variables that are readily available in external files, such as demographic characteristics or employment histories, ill-intentioned users – henceforth called intruders – may be able to link records in the released data to records in external files, thereby compromising the agency’s promise of confidentiality to those who provided the data. In response to this threat, agencies have developed an impressive variety of strategies for reducing the risks of unintended disclosures, ranging from restricting data access to altering data before release. Strategies that fall into the latter category are known as statistical disclosure limitation (SDL) techniques. Most SDL techniques have been developed for data derived from probability surveys or censuses. Even in complete form, these data would not typically be thought of as big data, with respect to scale (numbers of cases and attributes), complexity of attribute types, or structure: most datasets are released, if not actually structured, as flat files.
评估基于错误分类的调查微观数据披露限制方法所提供的保护
DOI: 10.1214/09-aoas317
发表时间: 2010
期刊: The Annals of Applied Statistics
影响因子: --
作者:
Shlomo N
通讯作者: Shlomo N