Optimally splitting cases for training and testing high dimensional classifiers.

Optimally splitting cases for training and testing high dimensional classifiers.
复制标题

DOI:
10.1186/1755-8794-4-31
复制
发表时间:
2011-04-08
影响因子:
2.7
通讯作者:
Simon RM
Simon RM
中科院分区:
医学3区
文献类型:
--
作者:
Dobbin KK;Simon RM

文献摘要

参考文献

被引文献

相似文献

我们考虑设计一个从高维数据开发预测分类器的研究问题。一种常见的研究设计是将样本分成训练集和独立测试集,前者用于开发分类器,后者用于评估其性能。在本文中,我们解决了应该将多少比例的样本用于训练集的问题。这个比例如何影响预测精度估计的均方误差(MSE) ?我们开发了一种非参数算法,用于确定可用于特定数据集和分类器算法的最佳分割比例。我们还进行了广泛的模拟研究,以便更好地理解决定最佳分割比例的因素,并在各种条件下评估常用的分割策略(1/2训练或2/3训练)。这些方法基于将MSE分解为三个直观的组成部分。通过将这些方法应用于许多合成和真实的微阵列数据集,我们表明,对于线性分类器,最佳比例取决于可用样本的总数和类别之间的差异表达程度。发现最佳比例取决于完整的数据集大小(n)和分类精度——更高的精度和更小的n导致更多的分配给训练集。对于具有强信号的合理大小的数据集(n≥100)(即85%或更高的完整数据集精度),通常使用的分配2/3的案例用于训练的策略接近最优。一般来说,我们建议使用我们的非参数重采样方法来确定最佳分割。这种方法可以应用于任何数据集,使用任何预测器开发方法,以确定最佳分割。
We consider the problem of designing a study to develop a predictive classifier from high dimensional data. A common study design is to split the sample into a training set and an independent test set, where the former is used to develop the classifier and the latter to evaluate its performance. In this paper we address the question of what proportion of the samples should be devoted to the training set. How does this proportion impact the mean squared error (MSE) of the prediction accuracy estimate? We develop a non-parametric algorithm for determining an optimal splitting proportion that can be applied with a specific dataset and classifier algorithm. We also perform a broad simulation study for the purpose of better understanding the factors that determine the best split proportions and to evaluate commonly used splitting strategies (1/2 training or 2/3 training) under a wide variety of conditions. These methods are based on a decomposition of the MSE into three intuitive component parts. By applying these approaches to a number of synthetic and real microarray datasets we show that for linear classifiers the optimal proportion depends on the overall number of samples available and the degree of differential expression between the classes. The optimal proportion was found to depend on the full dataset size (n) and classification accuracy - with higher accuracy and smaller n resulting in more assigned to the training set. The commonly used strategy of allocating 2/3rd of cases for training was close to optimal for reasonable sized datasets (n ≥ 100) with strong signals (i.e. 85% or greater full dataset accuracy). In general, we recommend use of our nonparametric resampling approach for determing the optimal split. This approach can be applied to any dataset, using any predictor development method, to determine the best split.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1093/bioinformatics/bth461
发表时间: 2005-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Fu, WJJ;Dougherty, ER;Carroll, RJ
通讯作者: Carroll, RJ
DOI: 10.1101/gr.184501
发表时间: 2001-11-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Boer, JM;Huber, WK;Poustka, A
通讯作者: Poustka, A
DOI: 10.1093/bioinformatics/bti499
发表时间: 2005-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Molinaro, AM;Simon, R;Pfeiffer, RM
通讯作者: Pfeiffer, RM
DOI: 10.1016/s0047-259x(03)00096-4
发表时间: 2004-02-01
影响因子: 1.6
作者:
Ledoit, O;Wolf, M
通讯作者: Wolf, M