Gaining Outlier Resistance With Progressive Quantiles: Fast Algorithms and Theoretical Studies

Gaining Outlier Resistance With Progressive Quantiles: Fast Algorithms and Theoretical Studies
复制标题

DOI:
10.1080/01621459.2020.1850460
复制
发表时间:
2020-11
影响因子:
3.7
通讯作者:
Yiyuan She;Zhifeng Wang;Jiahui Shen
Yiyuan She;Zhifeng Wang;Jiahui Shen
中科院分区:
数学1区
文献类型:
--
作者:
Yiyuan She;Zhifeng Wang;Jiahui Shen

文献摘要

被引文献

相似文献

摘要离群点广泛存在于大数据应用中,并可能严重影响统计估计和推理。本文引入了一种抗异常值估计的框架,以使任意给定的损失函数具有鲁棒性。它与修剪方法有密切的联系,并包括所有样本的显式突出度参数,这反过来又便于计算、理论和参数调整。为了解决非凸性和非光滑性问题,我们开发了易于实现并保证快速收敛的可伸缩算法。特别是,提出了一种新的技术来降低对起始点的要求,使得在规则的数据集上,数据重采样的次数可以大大减少。基于统计和计算相结合的处理,我们能够进行超出M-估计的非渐近分析。所得到的抵抗估计虽然不一定是全局最优的,甚至不一定是局部最优的,但在低维和高维上都具有极小极大似然最优性。在回归、分类和神经网络中的实验表明,所提出的方法在出现粗大离群点时具有很好的性能。这篇文章的补充材料可以在网上找到。
Abstract Outliers widely occur in big-data applications and may severely affect statistical estimation and inference. In this article, a framework of outlier-resistant estimation is introduced to robustify an arbitrarily given loss function. It has a close connection to the method of trimming and includes explicit outlyingness parameters for all samples, which in turn facilitates computation, theory, and parameter tuning. To tackle the issues of nonconvexity and nonsmoothness, we develop scalable algorithms with implementation ease and guaranteed fast convergence. In particular, a new technique is proposed to alleviate the requirement on the starting point such that on regular datasets, the number of data resamplings can be substantially reduced. Based on combined statistical and computational treatments, we are able to perform nonasymptotic analysis beyond M-estimation. The obtained resistant estimators, though not necessarily globally or even locally optimal, enjoy minimax rate optimality in both low dimensions and high dimensions. Experiments in regression, classification, and neural networks show excellent performance of the proposed methodology at the occurrence of gross outliers. Supplementary materials for this article are available online.