Toward Generalizing the Unification with Statistical Outliers

Toward Generalizing the Unification with Statistical Outliers
复制标题

DOI:
10.1145/2829956
复制
发表时间:
2016-01
期刊:
ACM Transactions on Knowledge Discovery from Data (TKDD)
影响因子:
--
通讯作者:
F. Angiulli;Fabio Fassetti
F. Angiulli;Fabio Fassetti
中科院分区:
其他
文献类型:
--
作者:
F. Angiulli;Fabio Fassetti

文献摘要

被引文献

相似文献

在这项工作中,我们引入了一种新颖的异常值定义,即梯度异常值因子(或 GOF),旨在提供一种与某些标准分布上的统计定义相统一的定义,但在存在混合分布的情况下具有不同的行为。直观上,GOF 分数衡量的是留在某个物体附近的概率。它与密度成正比,与密度的变化成反比。我们推导了 GOF 定义统一统计异常值定义的形式属性,并表明这种统一适用于某些标准分布,而 GOF 能够在存在不同分布的情况下捕获尾部,即使它们的密度明显不同。此外,我们通过数据密度的密度概念提供了 GOF 分数的概率解释。实验结果证实,在某些情况下,新的定义可以被有效利用。据我们所知,除了基于距离的异常值之外,没有其他数据挖掘异常值定义与统计异常值具有如此清晰的关系。
In this work, we introduce a novel definition of outlier, namely the Gradient Outlier Factor (or GOF), with the aim to provide a definition that unifies with the statistical one on some standard distributions but has a different behavior in the presence of mixture distributions. Intuitively, the GOF score measures the probability to stay in the neighborhood of a certain object. It is directly proportional to the density and inversely proportional to the variation of the density. We derive formal properties under which the GOF definition unifies the statistical outlier definition and show that the unification holds for some standard distributions, while the GOF is able to capture tails in the presence of different distributions even if their densities sensibly differ. Moreover, we provide a probabilistic interpretation of the GOF score, by means of the notion of density of the data density. Experimental results confirm that there are scenarios in which the novel definition can be profitably employed. To the best of our knowledge, except for distance-based outlier, no other data mining outlier definition has a so clearly established relationship with statistical outliers.