Selection bias in gene extraction on the basis of microarray gene-expression data

Selection bias in gene extraction on the basis of microarray gene-expression data
复制标题

DOI:
10.1073/pnas.102102699
复制
发表时间:
2002-05-14
影响因子:
11.1
通讯作者:
McLachlan, GJ
McLachlan, GJ
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Ambroise, C;McLachlan, GJ

文献摘要

被引文献

相似文献

在癌症诊断和治疗的背景下,我们考虑基于相对少量的已知类型的肿瘤组织样本构建准确的预测规则的问题,其中包含非常多(可能数千个)基因的表达数据。最近,文献中提出的结果表明,可以仅从少数基因构建预测规则,从而使其预测错误率可以忽略不计。然而,在这些结果中,测试误差或留一交叉验证误差的计算没有考虑选择偏差。没有允许,因为该规则要么在首先用于选择规则中使用的基因的组织样本上进行测试,要么因为规则的交叉验证不在选择过程之外;也就是说,在交叉验证过程的每个阶段训练规则时不进行基因选择。我们描述了在实践中如何通过执行交叉验证或在选择过程外部应用引导程序来评估和纠正选择偏差。我们建议使用 10 折交叉验证而不是留一交叉验证,并且关于引导程序,我们建议使用所谓的。 632+ 引导误差估计,旨在处理过度拟合的预测规则。使用两个已发表的数据集,我们证明,当对选择偏差进行校正时,对于仅少数基因的子集,交叉验证的误差不再为零。
In the context of cancer diagnosis and treatment, we consider the problem of constructing an accurate prediction rule on the basis of a relatively small number of tumor tissue samples of known type containing the expression data on very many (possibly thousands) genes. Recently, results have been presented in the literature suggesting that it is possible to construct a prediction rule from only a few genes such that it has a negligible prediction error rate. However, in these results the test error or the leave-one-out cross-validated error is calculated without allowance for the selection bias. There is no allowance because the rule is either tested on tissue samples that were used in the first instance to select the genes being used in the rule or because the cross-validation of the rule is not external to the selection process; that is, gene selection is not performed in training the rule at each stage of the cross-validation process. We describe how in practice the selection bias can be assessed and corrected for by either performing a cross-validation or applying the bootstrap external to the selection process. We recommend using 10-fold rather than leave-one-out cross-validation, and concerning the bootstrap, we suggest using the so-called. 632+ bootstrap error estimate designed to handle overfitted prediction rules. Using two published data sets, we demonstrate that when correction is made for the selection bias, the cross-validated error is no longer zero for a subset of only a few genes.