Improved method for correcting sample Mahalanobis distance without estimating population eigenvalues or eigenvectors of covariance matrix

Improved method for correcting sample Mahalanobis distance without estimating population eigenvalues or eigenvectors of covariance matrix
复制标题

无需估计协方差矩阵总体特征值或特征向量的修正样本马氏距离的改进方法

DOI:
10.1007/s41060-019-00201-4
复制
发表时间:
2019
期刊:
International Journal of Data Science and Analysis
影响因子:
--
通讯作者:
Yasuyuki Kobayashi
Yasuyuki Kobayashi
中科院分区:
--
文献类型:
--
作者:
Yasuyuki Kobayashi

文献摘要

被引文献

相似文献

样本马氏距离(SMD)的识别性能随着学习样本数量的减少而变差。因此,校正总体马哈拉诺比斯距离 (PMD) 的 SMD 非常重要,使其等效于无限学习样本的情况。为了减少这一主要目的的计算时间和成本,本文提出了一种不需要估计协方差矩阵的总体特征值或特征向量的校正方法。简而言之,该方法只需要协方差矩阵的样本特征值、学习样本数和维数即可对 PMD 进行 SMD 校正。该方法涉及 SMD 主成分的求和(每个主成分除以使用 delta 方法获得的期望)、Lawley 偏差估计以及样本特征向量的方差。数值实验表明,该方法对于学习样本数、维数、总体特征值序列和非中心性的各种情况都有效。该方法的应用还表明使用期望最大化算法估计高斯混合模型的性能得到了改善。
The recognition performance of the sample Mahalanobis distance (SMD) deteriorates as the number of learning samples decreases. Therefore, it is important to correct the SMD for a population Mahalanobis distance (PMD) such that it becomes equivalent to the case of infinite learning samples. In order to reduce the computation time and cost for this main purpose, this paper presents a correction method that does not require the estimation of the population eigenvalues or eigenvectors of the covariance matrix. In short, this method only requires the sample eigenvalues of the covariance matrix, number of learning samples, and dimensionality to correct the SMD for the PMD. This method involves the summation of the SMD’s principal components (each of which is divided by its expectation obtained using the delta method), Lawley’s bias estimation, and the variances of the sample eigenvectors. A numerical experiment demonstrates that this method works well for various cases of learning sample number, dimensionality, population eigenvalues sequence, and non-centrality. The application of this method also shows improved performance of estimating a Gaussian mixture model using the expectation–maximization algorithm.