The Power of Linear Reconstruction Attacks

The Power of Linear Reconstruction Attacks
复制标题

线性重建攻击的威力

DOI:
10.1137/1.9781611973105.102
复制
发表时间:
2012
期刊:
ArXiv
影响因子:
--
通讯作者:
Adam D. Smith
Adam D. Smith
中科院分区:
--
文献类型:
--
作者:
S. Kasiviswanathan;M. Rudelson;Adam D. Smith

文献摘要

被引文献

相似文献

我们考虑了线性重建攻击在统计数据隐私中的作用,表明它们可以应用于比以前理解的更广泛的设置。线性攻击之前已经被研究过(Dinur和Nissim PODS'03,Dwork,McSherry和Talwar STOC'07,Kasiviswanathan,Rudelson,Smith和Ullman STOC'10,De TCC'12,Muthukrishnan和Nikolov STOC'12),但到目前为止只应用于明显线性释放的设置。 考虑一个数据库管理员,他管理敏感信息的数据库,但希望发布关于数据库中的敏感属性(比如疾病)如何与一些非敏感属性(例如,邮政编码、年龄、性别等)。我们可以安装线性重建攻击的基础上,任何释放,使:a)满足给定的非退化布尔函数的记录的分数。这些释放包括列联表(Kasiviswanathan等人先前研究过,STOC'10)以及更复杂的输出,例如决策树等分类器的错误率; B)一大类M估计量中的任何一个(即经验风险最小化算法的输出),包括线性和逻辑回归的标准估计量。 我们做两个贡献:首先,我们展示了如何将这些类型的释放转换成线性格式,使它们服从于现有的多项式时间重建算法。这可能已经令人惊讶了,因为上面的许多版本(如M-估计)是通过求解高度非线性公式获得的。其次,我们展示了如何分析各种分布假设下的数据所产生的攻击。具体来说,我们考虑一个设置,其中发布了关于敏感属性如何与大小为k的所有子集(总共d个)非敏感布尔属性相关的相同统计量(上面的a)或B)。
We consider the power of linear reconstruction attacks in statistical data privacy, showing that they can be applied to a much wider range of settings than previously understood. Linear attacks have been studied before (Dinur and Nissim PODS'03, Dwork, McSherry and Talwar STOC'07, Kasiviswanathan, Rudelson, Smith and Ullman STOC'10, De TCC'12, Muthukrishnan and Nikolov STOC'12) but have so far been applied only in settings with releases that are obviously linear. Consider a database curator who manages a database of sensitive information but wants to release statistics about how a sensitive attribute (say, disease) in the database relates to some nonsensitive attributes (e.g., postal code, age, gender, etc). We show one can mount linear reconstruction attacks based on any release that gives: a) the fraction of records that satisfy a given non-degenerate boolean function. Such releases include contingency tables (previously studied by Kasiviswanathan et al., STOC'10) as well as more complex outputs like the error rate of classifiers such as decision trees; b) any one of a large class of M-estimators (that is, the output of empirical risk minimization algorithms), including the standard estimators for linear and logistic regression. We make two contributions: first, we show how these types of releases can be transformed into a linear format, making them amenable to existing polynomial-time reconstruction algorithms. This is already perhaps surprising, since many of the above releases (like M-estimators) are obtained by solving highly nonlinear formulations. Second, we show how to analyze the resulting attacks under various distributional assumptions on the data. Specifically, we consider a setting in which the same statistic (either a) or b) above) is released about how the sensitive attribute relates to all subsets of size k (out of a total of d) nonsensitive boolean attributes.