High-dimensional and large-scale anomaly detection using a linear one-class SVM with deep learning

High-dimensional and large-scale anomaly detection using a linear one-class SVM with deep learning
复制标题

DOI:
10.1016/j.patcog.2016.03.028
复制
发表时间:
2016-10-01
影响因子:
8
通讯作者:
Leckie, Christopher
Leckie, Christopher
中科院分区:
计算机科学1区
文献类型:
--
作者:
Erfani, Sarah M.;Rajasegarar, Sutharshan;Leckie, Christopher

文献摘要

被引文献

相似文献

高维问题域对异常检测提出了重大挑战。不相关特征的存在可以掩盖异常的存在。这个问题,被称为“维数灾难”,是许多异常检测技术的障碍。构建用于高维空间的鲁棒异常检测模型需要结合无监督特征提取器和异常检测器。虽然单类支持向量机在从行为良好的特征向量生成决策表面方面是有效的,但它们在对大型高维数据集的变化建模时可能效率低下。诸如深度信念网络(DBN)之类的架构是用于学习鲁棒特征的有前途的技术。我们提出了一个混合模型,其中一个无监督的DBN训练提取通用的底层功能,并从DBN学习的功能训练一类SVM。由于在我们的混合模型中,线性核可以代替非线性核而不损失精度,因此我们的模型是可扩展的,计算效率高。实验结果表明,我们提出的模型与深度自动编码器的异常检测性能相当,同时将其训练和测试时间分别减少了3倍和1000倍。(C)2016爱思唯尔有限公司版权所有
High-dimensional problem domains pose significant challenges for anomaly detection. The presence of irrelevant features can conceal the presence of anomalies. This problem, known as the 'curse of dimensionality', is an obstacle for many anomaly detection techniques. Building a robust anomaly detection model for use in high-dimensional spaces requires the combination of an unsupervised feature extractor and an anomaly detector. While one-class support vector machines are effective at producing decision surfaces from well-behaved feature vectors, they can be inefficient at modelling the variation in large, high-dimensional datasets. Architectures such as deep belief networks (DBNs) are a promising technique for learning robust features. We present a hybrid model where an unsupervised DBN is trained to extract generic underlying features, and a one-class SVM is trained from the features learned by the DBN. Since a linear kernel can be substituted for nonlinear ones in our hybrid model without loss of accuracy, our model is scalable and computationally efficient. The experimental results show that our proposed model yields comparable anomaly detection performance with a deep autoencoder, while reducing its training and testing time by a factor of 3 and 1000, respectively. (C) 2016 Elsevier Ltd. All rights reserved.