Multiple Distribution Data Description Learning Algorithm for Novelty Detection

Multiple Distribution Data Description Learning Algorithm for Novelty Detection
复制标题

用于新颖性检测的多分布数据描述学习算法

DOI:
10.1007/978-3-642-20847-8_21
复制
发表时间:
2011
影响因子:
5.4
通讯作者:
D. Sharma
D. Sharma
中科院分区:
工程技术1区
文献类型:
--
作者:
Trung Le;D. Tran;Wanli Ma;D. Sharma

文献摘要

参考文献

被引文献

相似文献

当前用于新奇检测的数据描述学习方法,例如支持向量数据描述和具有大余量的小球体,在正常数据集周围构造球形边界,以将该数据集与异常数据分离。该球体的体积被最小化以减少接受异常数据的机会。然而,这些学习方法并不能保证单一的球形边界可以最好地描述正常的数据集,如果这个集合中存在一些独特的数据分布。在本文中,我们提出了一种新的数据描述学习方法,构造一组球形的边界,以提供一个更好的数据描述正常的数据集。提出了一个优化问题,解决这个问题的结果在一个迭代学习算法,以确定一组球形边界。我们证明了在我们的学习方法中,每次迭代后的分类错误将减少。在28个已知数据集上的实验结果表明,该方法具有较低的分类错误率。
Current data description learning methods for novelty detection such as support vector data description and small sphere with large margin construct a spherically shaped boundary around a normal data set to separate this set from abnormal data. The volume of this sphere is minimized to reduce the chance of accepting abnormal data. However those learning methods do not guarantee that the single spherically shaped boundary can best describe the normal data set if there exist some distinctive data distributions in this set. We propose in this paper a new data description learning method that constructs a set of spherically shaped boundaries to provide a better data description to the normal data set. An optimisation problem is proposed and solving this problem results in an iterative learning algorithm to determine the set of spherically shaped boundaries. We prove that the classification error will be reduced after each iteration in our learning method. Experimental results on 28 well-known data sets show that the proposed method provides lower classification error rates.
DOI: --
发表时间: 2006-12
期刊: J. Mach. Learn. Res.
影响因子: --
作者:
Régis Vert;Jean-Philippe Vert
通讯作者: Régis Vert;Jean-Philippe Vert