Estimating the probabilities of misclassification using CV when the dimension and the sample sizes are large

Estimating the probabilities of misclassification using CV when the dimension and the sample sizes are large
复制标题

当维度和样本量很大时,使用 CV 估计误分类的概率

DOI:
10.32917/hmj/1544238034
复制
发表时间:
2018
影响因子:
0.2
通讯作者:
Tomoyuki Nakagawa
Tomoyuki Nakagawa
中科院分区:
数学4区
文献类型:
--
作者:
Tomoyuki Nakagawa

文献摘要

被引文献

相似文献

在本文中,我们研究估计高维数据中错误分类的概率。在许多情况下,交叉验证(CV)通常用于估计错误分类的概率。当样本量很大时,CV 使用原始数据提供近乎无偏的估计。另一方面,当维度比样本量大时,CV 的特性并不为人所知。因此,我们研究当维度和样本量趋于大时 CV 的渐近性质。此外,我们建议使用可用于高维数据的 CV 来纠正偏差的三种方法。我们展示了模拟研究中估计器的性能。
In this paper, we study about estimating the probabilities of misclassification in the high-dimensional data. In many cases, the cross-validation (CV) is often used for estimations of the probabilities of misclassification. CV provides a nearly unbiased estimate, using the original data when the sample sizes are large. On the other hand, the properties of CV are not well-known when the dimension is large as compared to the sample sizes. Therefore, we investigate asymptotic properties of CV when the dimension and the sample sizes tend to be large. Furthermore, we suggest the three methods for correcting the bias by using CV which is usable in the high-dimensional data. We show performances of the estimators in the simulation studies.