High-breakdown linear discriminant analysis

High-breakdown linear discriminant analysis
复制标题

DOI:
10.2307/2291457
复制
发表时间:
1997-03-01
影响因子:
3.7
通讯作者:
McLachlan, GJ
McLachlan, GJ
中科院分区:
数学1区
文献类型:
--
作者:
Hawkins, DM;McLachlan, GJ

文献摘要

被引文献

相似文献

线性判别分析的分类规则由数据来源的总体的真均值向量和共同协方差矩阵定义。因为这些真实参数通常是未知的,所以它们通常由从每个总体随机抽取的训练样本中的数据的样本均值向量和协方差矩阵估计。然而,这些样本统计非常容易受到离群值的污染,这一问题由于离群值可能对常规诊断不可见而变得更加复杂。高细分估计是一种程序,旨在消除这一令人担忧的原因,通过产生估计,免疫少数离群值的严重扭曲,无论其严重程度如何。在这篇文章中,我们激励和发展的线性判别分析的高故障标准,并给出其实现的算法。该程序的目的是补充,而不是取代通常的样本矩判别分析方法,无论是通过提供数据集不受离群值(支持通常的分析)的严重影响的迹象,或通过识别明显的异常点,并给出不受影响的抗性估计。
The classification rules of linear discriminant analysis are defined by the true mean vectors and the common covariance matrix of the populations from which the data come. Because these true parameters are generally unknown, they are commonly estimated by the sample mean vector and covariance matrix of the data in a training sample randomly drawn from each population. However, these sample statistics are notoriously susceptible to contamination by outliers, a problem compounded by the fact that the outliers may be invisible to conventional diagnostics. High-breakdown estimation is a procedure designed to remove this cause for concern by producing estimates that are immune to serious distortion by a minority of outliers, regardless of their severity. In this article we motivate and develop a high-breakdown criterion for linear discriminant analysis and give an algorithm for its implementation. The procedure is intended to supplement rather than replace the usual sample-moment methodology of discriminant analysis either by providing indications that the dataset is not seriously affected by outliers (supporting the usual analysis) or by identifying apparently aberrant points and giving resistant estimators that are not affected by them.