The Tao of Inference in Privacy-Protected Databases

The Tao of Inference in Privacy-Protected Databases
复制标题

DOI:
10.14778/3236187.3236217
复制
发表时间:
2018-07
期刊:
IACR Cryptol. ePrint Arch.
影响因子:
--
通讯作者:
Vincent Bindschaedler;Paul Grubbs;David Cash;Thomas Ristenpart;Vitaly Shmatikov
Vincent Bindschaedler;Paul Grubbs;David Cash;Thomas Ristenpart;Vitaly Shmatikov
中科院分区:
其他
文献类型:
--
作者:
Vincent Bindschaedler;Paul Grubbs;David Cash;Thomas Ristenpart;Vitaly Shmatikov

文献摘要

被引文献

相似文献

为了在支持标准功能的同时保护数据库的机密性,即使在面临完全妥协的情况下,最近的学术提案和商业产品都依赖于加密方案的混合。建议对“敏感”列应用强大的、语义上安全的加密,并使用支持排序等操作的属性揭示加密(PRE)来保护其他列。我们设计,实现和评估一种新的方法来推断存储在这样的加密数据库中的数据。其基石是多项式攻击,这是一种新的推理技术,在分析上是最优的,并且在经验上优于针对PRE加密数据的先前启发式攻击。我们还扩展了多项式攻击,以利用多列之间的相关性。这将以足够的准确度恢复预加密的数据,然后应用机器学习和记录链接方法来推断受语义安全加密或编辑保护的列。我们评估我们的方法对医疗,人口普查和工会会员数据集,首次展示如何推断完整的数据库记录。对于人口统计和邮政编码等预加密属性,我们的攻击比最佳先验启发式算法的性能高出16倍。与任何先前的技术不同,我们还推断出受强加密保护的属性,例如收入和医疗诊断。例如,当我们推断出出院数据集中的患者患有精神健康或药物滥用状况时,该预测的准确率为97%。
To protect database confidentiality even in the face of full compromise while supporting standard functionality, recent academic proposals and commercial products rely on a mix of encryption schemes. The recommendation is to apply strong, semantically secure encryption to the "sensitive" columns and protect other columns with property-revealing encryption (PRE) that supports operations such as sorting. We design, implement, and evaluate a new methodology for inferring data stored in such encrypted databases. The cornerstone is the multinomial attack , a new inference technique that is analytically optimal and empirically outperforms prior heuristic attacks against PRE-encrypted data. We also extend the multinomial attack to take advantage of correlations across multiple columns. This recovers PRE-encrypted data with sufficient accuracy to then apply machine learning and record linkage methods to infer columns protected by semantically secure encryption or redaction. We evaluate our methodology on medical, census, and union-membership datasets, showing for the first time how to infer full database records. For PRE-encrypted attributes such as demographics and ZIP codes, our attack outperforms the best prior heuristic by a factor of 16. Unlike any prior technique, we also infer attributes, such as incomes and medical diagnoses, protected by strong encryption. For example, when we infer that a patient in a hospital-discharge dataset has a mental health or substance abuse condition, this prediction is 97% accurate.