Estimating misclassification error with small samples via bootstrap cross-validation

Estimating misclassification error with small samples via bootstrap cross-validation
复制标题

DOI:
10.1093/bioinformatics/bti294
复制
发表时间:
2005-05-01
期刊:
影响因子:
5.8
通讯作者:
Wang, SJ
Wang, SJ
中科院分区:
生物学3区
文献类型:
--
作者:
Fu, WJJ;Carroll, RJ;Wang, SJ

文献摘要

被引文献

相似文献

动机:在临床诊断和生物信息学研究中,特别是在小样本微阵列研究中,误分类误差的估计受到越来越多的关注。现有的误差估计方法要么存在较大的变异性(如留一交叉验证),要么存在较大的偏倚(如重替换和留一自助法),因此不能令人满意。虽然小样本量仍然是昂贵的临床调查或微阵列研究,在资金,时间和组织材料的资源有限的关键特征之一,准确和易于实现的小样本误差估计方法是可取的,将是beneficial.Results:一个自举交叉验证方法进行了研究。它通过一个简单的过程与引导复位实现准确的误差估计,只有成本的计算机CPU时间。对微阵列数据的模拟研究和应用表明,它的性能始终优于其竞争对手。该方法具有几个吸引人的特性:(1)它是通过一个简单的程序实现;(2)它表现良好的小样本与样本大小,小到16;(3)它不限于任何特定的分类规则,因此适用于许多参数或非参数的方法。
Motivation: Estimation of misclassification error has received increasing attention in clinical diagnosis and bioinformatics studies, especially in small sample studies with microarray data. Current error estimation methods are not satisfactory because they either have large variability (such as leave-one-out cross-validation) or large bias (such as resubstitution and leave-one-out bootstrap). While small sample size remains one of the key features of costly clinical investigations or of microarray studies that have limited resources in funding, time and tissue materials, accurate and easy-to-implement error estimation methods for small samples are desirable and will be beneficial.Results: A bootstrap cross-validation method is studied. It achieves accurate error estimation through a simple procedure with bootstrap resampling and only costs computer CPU time. Simulation studies and applications to microarray data demonstrate that it performs consistently better than its competitors. This method possesses several attractive properties: (1) it is implemented through a simple procedure; (2) it performs well for small samples with sample size, as small as 16; (3) it is not restricted to any particular classification rules and thus applies to many parametric or non-parametric methods.