The performance of diagnostic-robust generalized potentials for the identification of multiple high leverage points in linear regression

The performance of diagnostic-robust generalized potentials for the identification of multiple high leverage points in linear regression
复制标题

DOI:
10.1080/02664760802553463
复制
发表时间:
2009-05
影响因子:
1.5
通讯作者:
M. Habshaha;M. R. Norazanb;Rahmatullah Imonc
M. Habshaha;M. R. Norazanb;Rahmatullah Imonc
中科院分区:
数学4区
文献类型:
--
作者:
M. Habshaha;M. R. Norazanb;Rahmatullah Imonc

文献摘要

被引文献

相似文献

在回归诊断中,杠杆值被用作X空间中有影响的观测值的度量。检测高杠杆值是至关重要的,因为它们负责对回归模型的拟合得出误导性结论,导致多重共线性问题,掩盖和/或淹没异常值等。在单个高杠杆点的识别方面已经做了大量的工作,一般认为单个高杠杆点的检测问题已经基本解决。但统计学家对多个高杠杆点的检测并没有达成普遍共识。当数据集中存在一组高杠杆点时,主要是因为屏蔽和/或淹没效应,常用的诊断方法无法正确识别它们。另一方面,稳健的替代方法可以正确识别高杠杆点,但它们往往会识别太多的低杠杆点而不是高杠杆点,这也是不希望的。人们试图在这两种方法之间作出妥协。我们提出了一种自适应方法,其中通过鲁棒方法识别可疑的高杠杆点,然后在诊断检查后将低杠杆点(如果有的话)放回估计数据集中。通过一些著名的数据集和蒙特卡罗模拟,研究了我们新提出的方法对多个高杠杆点检测的有效性。
Leverage values are being used in regression diagnostics as measures of influential observations in the $X$-space. Detection of high leverage values is crucial because of their responsibility for misleading conclusion about the fitting of a regression model, causing multicollinearity problems, masking and/or swamping of outliers, etc. Much work has been done on the identification of single high leverage points and it is generally believed that the problem of detection of a single high leverage point has been largely resolved. But there is no general agreement among the statisticians about the detection of multiple high leverage points. When a group of high leverage points is present in a data set, mainly because of the masking and/or swamping effects the commonly used diagnostic methods fail to identify them correctly. On the other hand, the robust alternative methods can identify the high leverage points correctly but they have a tendency to identify too many low leverage points to be points of high leverages which is not also desired. An attempt has been made to make a compromise between these two approaches. We propose an adaptive method where the suspected high leverage points are identified by robust methods and then the low leverage points (if any) are put back into the estimation data set after diagnostic checking. The usefulness of our newly proposed method for the detection of multiple high leverage points is studied by some well-known data sets and Monte Carlo simulations.