Provable De-anonymization of Large Datasets with Sparse Dimensions

Provable De-anonymization of Large Datasets with Sparse Dimensions
复制标题

稀疏维度大型数据集的可证明去匿名化

DOI:
--
复制
发表时间:
2012
期刊:
The post
影响因子:
--
通讯作者:
Arunesh Sinha
Arunesh Sinha
中科院分区:
--
文献类型:
--
作者:
Anupam Datta;Divya Sharma;Arunesh Sinha

文献摘要

被引文献

相似文献

有大量关于对包含关于个人的微观数据的数据库进行统计去匿名化攻击的经验性工作,例如,他们的偏好、电影评级或交易数据。我们的目标是分析解释为什么这种攻击工作。具体来说,我们分析了Narayanan-Shmatikov算法的一个变体,该算法用于有效地对Netflix电影评级数据库进行去匿名化。我们证明定理表征的数学性质的数据库和辅助信息提供给对手,使两类隐私攻击。在第一次攻击中,对手成功地识别出她拥有辅助信息的个人(孤立攻击)。在第二次攻击中,对手了解到关于个人的更多信息,尽管她可能无法唯一地识别他(信息放大攻击)。我们证明了分析结果的适用性,通过经验验证,假设的数据库的数学属性实际上是真实的Netflix电影评级数据库中的记录,其中包含约50万用户的评级的一个重要部分。
There is a significant body of empirical work on statistical de-anonymization attacks against databases containing micro-data about individuals, e.g., their preferences, movie ratings, or transaction data. Our goal is to analytically explain why such attacks work. Specifically, we analyze a variant of the Narayanan-Shmatikov algorithm that was used to effectively de-anonymize the Netflix database of movie ratings. We prove theorems characterizing mathematical properties of the database and the auxiliary information available to the adversary that enable two classes of privacy attacks. In the first attack, the adversary successfully identifies the individual about whom she possesses auxiliary information (an isolation attack). In the second attack, the adversary learns additional information about the individual, although she may not be able to uniquely identify him (an information amplification attack). We demonstrate the applicability of the analytical results by empirically verifying that the mathematical properties assumed of the database are actually true for a significant fraction of the records in the Netflix movie ratings database, which contains ratings from about 500,000 users.