How to Accurately and Privately Identify Anomalies

How to Accurately and Privately Identify Anomalies
复制标题

DOI:
10.1145/3319535.3363209
复制
发表时间:
2019-11
期刊:
Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
H. Asif;Periklis A. Papakonstantinou;Jaideep Vaidya
H. Asif;Periklis A. Papakonstantinou;Jaideep Vaidya
中科院分区:
其他
文献类型:
--
作者:
H. Asif;Periklis A. Papakonstantinou;Jaideep Vaidya

文献摘要

被引文献

相似文献

识别数据中的异常对于科学、国家安全和金融的进步至关重要。然而,隐私问题限制了我们分析数据的能力。我们能否取消这些限制,准确识别异常情况,而不损害贡献数据的人的隐私?我们在最实际相关的情况下解决这个问题,在这种情况下,一个记录被认为是相对于其他记录的异常。我们有四点贡献。首先,我们引入了敏感隐私的概念,它将私下识别异常的含义概念化。敏感隐私是区别隐私这一重要概念的概括,是可以分析的。重要的是,敏感隐私允许算法构造,提供强大且实际有意义的隐私和效用保证。其次,我们证明了差异隐私本质上不能准确地私下识别异常;从这个意义上说,我们的推广是必要的。第三,我们提供了一个通用编译器,它接受一个差异私有机制作为输入(该机制对异常识别的实用性很差),并将其转换为敏感私有机制。该编译器在理论上具有重要意义,它输出的机制的效用大大高于输入机制的效用。作为我们的第四个贡献,我们为异常((β,r)-异常)的流行定义提出了机制,该机制(I)被保证是敏感的私有的,(Ii)具有可证明的效用保证,以及(Iii)被经验证明在一系列数据集和评估标准上具有压倒性的准确性能。
Identifying anomalies in data is central to the advancement of science, national security, and finance. However, privacy concerns restrict our ability to analyze data. Can we lift these restrictions and accurately identify anomalies without hurting the privacy of those who contribute their data? We address this question for the most practically relevant case, where a record is considered anomalous relative to other records. We make four contributions. First, we introduce the notion of sensitive privacy, which conceptualizes what it means to privately identify anomalies. Sensitive privacy generalizes the important concept of differential privacy and is amenable to analysis. Importantly, sensitive privacy admits algorithmic constructions that provide strong and practically meaningful privacy and utility guarantees. Second, we show that differential privacy is inherently incapable of accurately and privately identifying anomalies; in this sense, our generalization is necessary. Third, we provide a general compiler that takes as input a differentially private mechanism (which has bad utility for anomaly identification) and transforms it into a sensitively private one. This compiler, which is mostly of theoretical importance, is shown to output a mechanism whose utility greatly improves over the utility of the input mechanism. As our fourth contribution we propose mechanisms for a popular definition of anomaly ((β,r)-anomaly) that (i) are guaranteed to be sensitively private, (ii) come with provable utility guarantees, and (iii) are empirically shown to have an overwhelmingly accurate performance over a range of datasets and evaluation criteria.