Increasing Adversarial Uncertainty to Scale Private Similarity Testing

Increasing Adversarial Uncertainty to Scale Private Similarity Testing
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Yiqing Hua;Armin Namavari;Kai-Wen Cheng;Mor Naaman;Thomas Ristenpart
Yiqing Hua;Armin Namavari;Kai-Wen Cheng;Mor Naaman;Thomas Ristenpart
中科院分区:
其他
文献类型:
--
作者:
Yiqing Hua;Armin Namavari;Kai-Wen Cheng;Mor Naaman;Thomas Ristenpart

文献摘要

相似文献

社交媒体和其他平台依赖于对滥用内容的自动检测,以帮助打击虚假信息、骚扰和滥用。一种常见的方法是根据服务器端有问题的项目数据库检查用户内容的相似性。然而,这种方法从根本上危及用户隐私。相反,我们的目标是客户端检测,当发生此类匹配时只通知用户,警告他们不要滥用内容。我们的解决方案是基于隐私保护的相似性测试。现有的方法依赖于昂贵的密码协议,这些协议不能很好地扩展到大型数据库,并且可能会牺牲匹配的正确性。为了应对这一挑战,我们提出并形式化了基于相似性的分段化的概念~(SBB)。使用SBB,客户端向数据库保存服务器显示少量信息,以便它可以生成一系列潜在的相似项。桶足够小,可以有效地应用基于相似性的隐私保护协议。为了分析泄露的信息的隐私风险,我们引入了一个框架来衡量对手在推断关于客户输入的正确谓词时的置信度。我们开发了一个实用的图像内容SBB协议,并用真实的社交媒体数据对其客户隐私保障进行了评估。然后,我们将SBB与各种相似性协议相结合,表明在大规模数据库上,与不使用SBB相比,使用SBB的速度至少提高了29倍,同时保持了95%以上的正确率。
Social media and other platforms rely on automated detection of abusive content to help combat disinformation, harassment, and abuse. One common approach is to check user content for similarity against a server-side database of problematic items. However, this method fundamentally endangers user privacy. Instead, we target client-side detection, notifying only the users when such matches occur to warn them against abusive content. Our solution is based on privacy-preserving similarity testing. Existing approaches rely on expensive cryptographic protocols that do not scale well to large databases and may sacrifice the correctness of the matching. To contend with this challenge, we propose and formalize the concept of similarity-based bucketization~(SBB). With SBB, a client reveals a small amount of information to a database-holding server so that it can generate a bucket of potentially similar items. The bucket is small enough for efficient application of privacy-preserving protocols for similarity. To analyze the privacy risk of the revealed information, we introduce a framework for measuring an adversary's confidence in inferring a predicate about the client input correctly. We develop a practical SBB protocol for image content, and evaluate its client privacy guarantee with real-world social media data. We then combine SBB with various similarity protocols, showing that the combination with SBB provides a speedup of at least 29x on large-scale databases compared to that without, while retaining correctness of over 95%.