How to Evaluate the Quality of Unsupervised Anomaly Detection Algorithms?

How to Evaluate the Quality of Unsupervised Anomaly Detection Algorithms?
复制标题

DOI:
--
复制
发表时间:
2016-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Nicolas Goix
Nicolas Goix
中科院分区:
其他
文献类型:
--
作者:
Nicolas Goix

文献摘要

被引文献

相似文献

当有足够的标记数据可用时,可以使用基于受试者操作特征(ROC)或精确-召回(PR)曲线的经典标准来比较无监督异常检测算法的性能。然而,在许多情况下,很少或没有数据被标记。这就需要人们可以在非标记数据上计算的替代标准。在本文中,两个标准,不需要标签的经验表明,以区分准确(w.r.t.基于ROC或PR的标准)。这些标准是基于现有的过剩质量(EM)和质量体积(MV)曲线,这通常不能很好地估计在大尺寸。还描述和测试了基于特征子采样和聚合的方法,将这些标准的使用扩展到高维数据集,并解决了标准EM和MV曲线固有的主要缺点。
When sufficient labeled data are available, classical criteria based on Receiver Operating Characteristic (ROC) or Precision-Recall (PR) curves can be used to compare the performance of un-supervised anomaly detection algorithms. However , in many situations, few or no data are labeled. This calls for alternative criteria one can compute on non-labeled data. In this paper, two criteria that do not require labels are empirically shown to discriminate accurately (w.r.t. ROC or PR based criteria) between algorithms. These criteria are based on existing Excess-Mass (EM) and Mass-Volume (MV) curves, which generally cannot be well estimated in large dimension. A methodology based on feature sub-sampling and aggregating is also described and tested, extending the use of these criteria to high-dimensional datasets and solving major drawbacks inherent to standard EM and MV curves.