Data-adaptive trimming of the Hill estimator and detection of outliers in the extremes of heavy-tailed data

Data-adaptive trimming of the Hill estimator and detection of outliers in the extremes of heavy-tailed data
复制标题

DOI:
10.1214/19-ejs1561
复制
发表时间:
2019-01-01
影响因子:
1.1
通讯作者:
Stoev, Stilian
Stoev, Stilian
中科院分区:
数学3区
文献类型:
--
作者:
Bhattacharya, Shrijita;Kallitsis, Michael;Stoev, Stilian

文献摘要

被引文献

相似文献

对于重尾分布的指数,我们引入了一个修剪形式的Hill估计,它对极值顺序统计量中的扰动是稳健的。在理想的Pareto环境下,在给定严格上断点的所有无偏估计中,估计量本质上是有限样本有效的。对于一般的重尾模型,我们在二阶正则变差条件下建立了估计量的渐近正态,并证明了它在Hall分布类中是极小极大速率最优的。我们还开发了一种自动的、数据驱动的方法来选择修剪参数,该方法产生了一种新型的稳健估计器,可以适应极端情况下的未知污染水平。这种自适应稳健性使得我们的估计器特别有吸引力,并且在数据极值被污染的情况下优于其他稳健估计器。作为数据驱动的剪裁参数选择的一个重要应用,我们得到了一种对重尾数据中的极端异常值进行原则性识别的方法。事实上,该方法已经被证明能够正确地识别先前探索的Condroz数据集中的离群值的数量。
We introduce a trimmed version of the Hill estimator for the index of a heavy-tailed distribution, which is robust to perturbations in the extreme order statistics. In the ideal Pareto setting, the estimator is essentially finite-sample efficient among all unbiased estimators with a given strict upper break-down point. For general heavy-tailed models, we establish the asymptotic normality of the estimator under second order regular variation conditions and also show that it is minimax rate-optimal in the Hall class of distributions. We also develop an automatic, data-driven method for the choice of the trimming parameter which yields a new type of robust estimator that can adapt to the unknown level of contamination in the extremes. This adaptive robustness property makes our estimator particularly appealing and superior to other robust estimators in the setting where the extremes of the data are contaminated. As an important application of the data-driven selection of the trimming parameters, we obtain a methodology for the principled identification of extreme outliers in heavy tailed data. Indeed, the method has been shown to correctly identify the number of outliers in the previously explored Condroz data set.