Multi-View Low-Rank Analysis with Applications to Outlier Detection

Multi-View Low-Rank Analysis with Applications to Outlier Detection
复制标题

DOI:
10.1145/3168363
复制
发表时间:
2018-03
期刊:
ACM Transactions on Knowledge Discovery from Data (TKDD)
影响因子:
--
通讯作者:
Sheng Li;Ming Shao;Y. Fu
Sheng Li;Ming Shao;Y. Fu
中科院分区:
其他
文献类型:
--
作者:
Sheng Li;Ming Shao;Y. Fu

文献摘要

被引文献

相似文献

检测异常值或离群点是各种机器学习和数据挖掘应用中的一个基本问题。传统的异常值检测算法主要是为单视角数据设计的。如今,可以很容易地从多个视角收集数据,并且许多学习任务,如聚类和分类,都受益于多视角数据。然而,从多视角数据中检测异常值仍然是一个极具挑战性的问题,因为多视角中的数据通常具有更复杂的分布,并表现出不一致的行为。为了解决这个问题,在本文中我们提出了一个用于异常值检测的多视角低秩分析(MLRA)框架。MLRA从一个新的视角——稳健的数据表示来寻找异常值。它包含两个主要部分。首先,进行跨视角低秩编码以揭示数据的内在结构。特别是,我们构建了一个正则化的秩最小化问题,并通过一种高效的优化算法来解决它。其次,通过一个异常值分数估计程序来识别异常值。与现有的多视角异常值检测方法不同,MLRA能够同时从多个视角检测两种不同类型的异常值。为此,我们设计了一个标准,通过分析所获得的表示系数来估计异常值分数。此外,我们扩展了MLRA以解决多视角组异常值检测问题。对七个UCI数据集、MovieLens数据集、USPS - MNIST数据集和WebKB数据集的大量评估表明,我们的方法优于几种最先进的异常值检测方法。
Detecting outliers or anomalies is a fundamental problem in various machine learning and data mining applications. Conventional outlier detection algorithms are mainly designed for single-view data. Nowadays, data can be easily collected from multiple views, and many learning tasks such as clustering and classification have benefited from multi-view data. However, outlier detection from multi-view data is still a very challenging problem, as the data in multiple views usually have more complicated distributions and exhibit inconsistent behaviors. To address this problem, we propose a multi-view low-rank analysis (MLRA) framework for outlier detection in this article. MLRA pursuits outliers from a new perspective, robust data representation. It contains two major components. First, the cross-view low-rank coding is performed to reveal the intrinsic structures of data. In particular, we formulate a regularized rank-minimization problem, which is solved by an efficient optimization algorithm. Second, the outliers are identified through an outlier score estimation procedure. Different from the existing multi-view outlier detection methods, MLRA is able to detect two different types of outliers from multiple views simultaneously. To this end, we design a criterion to estimate the outlier scores by analyzing the obtained representation coefficients. Moreover, we extend MLRA to tackle the multi-view group outlier detection problem. Extensive evaluations on seven UCI datasets, the MovieLens, the USPS-MNIST, and the WebKB datasets demon strate that our approach outperforms several state-of-the-art outlier detection methods.