Understanding the Effect of Bias in Deep Anomaly Detection

Understanding the Effect of Bias in Deep Anomaly Detection
复制标题

DOI:
10.24963/ijcai.2021/456
复制
发表时间:
2021-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Ziyu Ye;Yuxin Chen;Haitao Zheng
Ziyu Ye;Yuxin Chen;Haitao Zheng
中科院分区:
其他
文献类型:
--
作者:
Ziyu Ye;Yuxin Chen;Haitao Zheng

文献摘要

相似文献

由于标记为异常数据的稀缺性,异常检测在机器学习中提出了独特的挑战。最近的工作试图通过使用其他标记为异常样品的深度异常检测模型来减轻此类问题。但是,标记的数据通常与目标分布不符,并将有害偏差引入受过训练的模型。在本文中,我们旨在了解偏置异常对异常检测的影响。具体而言,我们将异常检测视为一项监督学习任务,目的是以给定的假阳性率优化召回率。我们正式研究了异常检测器的相对评分偏置,该探测器定义为相对于基线异常检测器的性能差异。我们建立了第一个有限的样本率,用于估计深度异常检测的相对评分偏差,并在经验上验证我们对合成和现实世界数据集的理论结果。我们还提供了一项广泛的实证研究,涉及偏置训练异常集如何影响异常得分函数,从而在不同异常类别上的检测性能。我们的研究表明,偏见的异常集可能有用或有问题,并为将来的研究提供了坚实的基准。
Anomaly detection presents a unique challenge in machine learning, due to the scarcity of labeled anomaly data. Recent work attempts to mitigate such problems by augmenting training of deep anomaly detection models with additional labeled anomaly samples. However, the labeled data often does not align with the target distribution and introduces harmful bias to the trained model. In this paper, we aim to understand the effect of a biased anomaly set on anomaly detection. Concretely, we view anomaly detection as a supervised learning task where the objective is to optimize the recall at a given false positive rate. We formally study the relative scoring bias of an anomaly detector, defined as the difference in performance with respect to a baseline anomaly detector. We establish the first finite sample rates for estimating the relative scoring bias for deep anomaly detection, and empirically validate our theoretical results on both synthetic and real-world datasets. We also provide an extensive empirical study on how a biased training anomaly set affects the anomaly score function and therefore the detection performance on different anomaly classes. Our study demonstrates scenarios in which the biased anomaly set can be useful or problematic, and provides a solid benchmark for future research.