Variable kernel density estimation based robust regression and its applications

Variable kernel density estimation based robust regression and its applications
复制标题

基于变核密度估计的鲁棒回归及其应用

DOI:
10.1016/j.neucom.2012.12.076
复制
发表时间:
2014-06
期刊:
影响因子:
6
通讯作者:
Yanning Zhang
Yanning Zhang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhen Zhang;Yanning Zhang

文献摘要

参考文献

相似文献

高故障点稳健估计是计算机视觉、机器学习等诸多领域的一个重要的基础性课题。传统的稳健性估计器的故障点超过50%,例如,随机抽样共识及其导数,需要一个用户指定的内点标度,以便能够区分内点和异常点,但在许多应用中,我们没有任何内点标度的先验知识,所以通常指定一个经验值。近年来,人们提出了一组基于核密度估计(KDE)的稳健估计来解决这一问题。然而,由于KDE最重要的参数带宽与内嵌节点的规模高度相关,这些方法被证明是内嵌节点规模的估计器,估计内嵌节点的规模并不是一件容易的事情。为此,作者构造了一种基于变核密度估计(VKDE)的稳健估计量。与KDE相比,VKDE利用样本的局部信息来估计带宽,使用K-近邻方法,而不是从节点的尺度估计带宽。因此,可以省略对内点规模的估计。此外,由于采用了变带宽技术,该方法在样本分布较密集的区域使用了较小的带宽。由于内值点的分布比离群点要密集得多,因此该方法对内值点的分辨率更高,估计密度的峰值将更接近样本分布最密集的点。最后将该方法与随机抽样一致性估计和最小中值平方估计两种最常用的稳健估计方法进行了比较。从结果可以看出,该方法比这两种方法具有更高的精度。
Robust estimation with high break down point is an important and fundamental topic in computer vision, machine learning and many other areas. Traditional robust estimator with a break down point more than 50%, for illustration, Random Sampling Consensus and its derivatives, needs a user specified scale of inliers such that inliers can be distinguished from outliers, but in many applications, we do not have any a priori of the scale of inliers, so an empirical value is usually specified. In recent years, a group of Kernel Density Estimation (KDE) based robust estimators has been proposed to solve this problem. However, as the most important parameter, bandwidth, for KDE is highly correlated to the scale of inliers, these methods turned out to be a scale estimator for inliers, and it is not an easy work to estimate the scale of inliers. Thus, the authors build up a robust estimator based on Variable Kernel Density Estimation (VKDE). Compared to KDE, VKDE estimates bandwidth out of local information of samples by usingK-Nearest-Neighbor method instead of estimating bandwidth from the scale of inliers. Thus the estimation for the scale of inliers can be omitted. Furthermore, as variable bandwidth technique is applied, the proposed method uses smaller bandwidths for the areas where samples are more densely distributed. As inliers are much more densely distributed than outliers, the proposed method achieved a higher resolution for inliers, and then the peak of estimated density will be closer to the point near which samples are most densely distributed. At last, the proposed method is compared to two most widely used robust estimators, Random Sampling Consensus and Least Median Square. From the result we can see that it has higher precision than those two methods.
DOI: 10.1109/tpami.2009.148
发表时间: 2010-01
影响因子: 23.6
作者:
Wang H;Mirota D;Hager GD
通讯作者: Hager GD
DOI: 10.1080/01621459.1996.10476701
发表时间: 1996-03
影响因子: 3.7
作者:
M. C. Jones;J. Marron;S. Sheather
通讯作者: M. C. Jones;J. Marron;S. Sheather
DOI: 10.1080/01621459.1984.10477105
发表时间: 1984-12
影响因子: 3.7
作者:
P. Rousseeuw
通讯作者: P. Rousseeuw
DOI: 10.1006/cviu.1999.0832
发表时间: 2000-04-01
影响因子: 4.5
作者:
Torr, PHS;Zisserman, A
通讯作者: Zisserman, A
DOI: 10.1017/cbo9780511811685.009
发表时间: 2001-04
期刊: --
影响因子: --
作者:
Richard Hartley;Andrew Zisserman
通讯作者: Richard Hartley;Andrew Zisserman