An information theoretic approach for privacy metrics

An information theoretic approach for privacy metrics
复制标题

DOI:
--
复制
发表时间:
2010-12
期刊:
Trans. Data Priv.
影响因子:
--
通讯作者:
M. Bezzi
M. Bezzi
中科院分区:
其他
文献类型:
--
作者:
M. Bezzi

文献摘要

被引文献

相似文献

组织通常需要在不泄露敏感信息的情况下发布微数据。在这个范围内,数据是匿名的,为了评估过程的质量,已经提出了各种隐私度量,如k-匿名性、L-多样性和t-闭合度。这些指标能够捕获披露风险的不同方面,对个人与敏感属性的关联施加最低要求。如果我们想要将它们结合到一个优化问题中,我们需要一个能够表达所有这些隐私条件的公共框架。以往的研究提出了互信息的概念来衡量不同类型的信息披露风险和效用,但由于互信息是一个平均值,不能将这些情况完整地表达在单一记录上。这里我们引入了一种符号信息的概念(即,单个记录对共同信息的贡献),它允许表达和比较披露风险度量。此外,我们还得到了风险值t与L之间的关系,该关系可用于参数设置。通过数值实验,我们还证明了L-多样性和t-贴近度如何用两个不同但同样可接受的信息增益条件来表示。
Organizations often need to release microdata without revealing sensitive information. To this scope, data are anonymized and, to assess the quality of the process, various privacy metrics have been proposed, such as k-anonymity, l-diversity, and t-closeness. These metrics are able to capture different aspects of the disclosure risk, imposing minimal requirements on the association of an individual with the sensitive attributes. If we want to combine them in a optimization problem, we need a common framework able to express all these privacy conditions. Previous studies proposed the notion of mutual information to measure the different kinds of disclosure risks and the utility, but, since mutual information is an average quantity, it is not able to completely express these conditions on single records. We introduce here the notion of one-symbol information (i.e., the contribution to mutual information by a single record) that allows to express and compare the disclosure risk metrics. In addition, we obtain a relation between the risk values t and l, which can be used for parameter setting. We also show, by numerical experiments, how l-diversity and t-closeness can be represented in terms of two different, but equally acceptable, conditions on the information gain..