Active Learning for Anomaly and Rare-Category Detection

Active Learning for Anomaly and Rare-Category Detection
复制标题

DOI:
--
复制
发表时间:
2004-12
期刊:
影响因子:
4.5
通讯作者:
D. Pelleg;A. Moore
D. Pelleg;A. Moore
中科院分区:
化学3区
文献类型:
--
作者:
D. Pelleg;A. Moore

文献摘要

被引文献

相似文献

我们介绍了一种新的主动学习的情况下,用户希望与学习算法来识别有用的异常。这些与异常的传统统计定义不同,即异常值或仅仅是模型不佳的点。我们的区别在于,异常的有用性是由用户主观分类的。我们做了两个额外的假设。首先,在一个庞大的数据集中,几乎没有什么有用的异常可以被发现。第二,有用和无用的异常有时可能存在于类似异常的微小类别中。因此,挑战是在人类专家的帮助下(以类别标签的形式)在未标记的噪声集中识别“稀有类别”记录,该人类专家具有他们准备分类的少量数据点预算。我们提出了一种技术,以满足这一挑战,它假设一个混合模型适合的数据,但在其他方面没有假设的特定形式的混合成分。该属性在现实生活场景和各种统计模型中具有广泛的适用性。我们给出了几种替代方法的概述,突出了它们的优点和缺点,并得出了详细的实证分析。我们证明了我们的方法可以快速放大包含数十万数据集中的几十个点的异常集。
We introduce a novel active-learning scenario in which a user wants to work with a learning algorithm to identify useful anomalies. These are distinguished from the traditional statistical definition of anomalies as outliers or merely ill-modeled points. Our distinction is that the usefulness of anomalies is categorized subjectively by the user. We make two additional assumptions. First, there exist extremely few useful anomalies to be hunted down within a massive dataset. Second, both useful and useless anomalies may sometimes exist within tiny classes of similar anomalies. The challenge is thus to identify "rare category" records in an unlabeled noisy set with help (in the form of class labels) from a human expert who has a small budget of datapoints that they are prepared to categorize. We propose a technique to meet this challenge, which assumes a mixture model fit to the data, but otherwise makes no assumptions on the particular form of the mixture components. This property promises wide applicability in real-life scenarios and for various statistical models. We give an overview of several alternative methods, highlighting their strengths and weaknesses, and conclude with a detailed empirical analysis. We show that our method can quickly zoom in on an anomaly set containing a few tens of points in a dataset of hundreds of thousands.