Optimal classifier selection and negative bias in error rate estimation: an empirical study on high-dimensional prediction

Optimal classifier selection and negative bias in error rate estimation: an empirical study on high-dimensional prediction
复制标题

DOI:
10.1186/1471-2288-9-85
复制
发表时间:
2009-12-21
影响因子:
4
通讯作者:
Strobl, Carolin
Strobl, Carolin
中科院分区:
医学3区
文献类型:
--
作者:
Boulesteix, Anne-Laure;Strobl, Carolin

文献摘要

被引文献

相似文献

背景:在生物识别实践中,研究人员经常在“试错”策略中应用大量不同的方法,以尽可能多地从数据中获取数据,并且由于出版压力或来自咨询客户的压力,只提供最有利的结果。这种策略可能会在预测误差估计中引起很大的乐观偏差,这在本手稿中进行了定量评估。我们工作的重点是基于高维数据(例如微阵列数据)的类别预测,因为此类分析特别容易受到这种偏差的影响。方法:在我们的研究中,我们在交叉验证评估方案中考虑了总共 124 个分类器变体(可能包括变量选择或调整步骤)。分类器应用于原始和修改后的真实微阵列数据集,其中一些分类器是通过随机排列类别标签来模拟非信息预测变量,同时保留其相关结构而获得的。结果:我们评估了分类器不同变体的最小错误分类率,以便量化以数据驱动方式事后选择最佳分类器时产生的偏差。分别和联合检查参数调整(包括作为特殊情况的基因选择参数)产生的偏差和分类方法选择产生的偏差。结论:基于结肠癌和前列腺癌研究中排列的无信息预测因子,所研究的分类器的中位最小错误率分别低至 31% 和 41%。我们的结论是,仅呈现最佳结果的策略是不可接受的,因为它会在错误率估计中产生很大的偏差,并建议正确报告分类准确性的替代方法。
Background: In biometric practice, researchers often apply a large number of different methods in a "trial-and-error" strategy to get as much as possible out of their data and, due to publication pressure or pressure from the consulting customer, present only the most favorable results. This strategy may induce a substantial optimistic bias in prediction error estimation, which is quantitatively assessed in the present manuscript. The focus of our work is on class prediction based on high-dimensional data (e. g. microarray data), since such analyses are particularly exposed to this kind of bias.Methods: In our study we consider a total of 124 variants of classifiers (possibly including variable selection or tuning steps) within a cross-validation evaluation scheme. The classifiers are applied to original and modified real microarray data sets, some of which are obtained by randomly permuting the class labels to mimic non-informative predictors while preserving their correlation structure.Results: We assess the minimal misclassification rate over the different variants of classifiers in order to quantify the bias arising when the optimal classifier is selected a posteriori in a data-driven manner. The bias resulting from the parameter tuning (including gene selection parameters as a special case) and the bias resulting from the choice of the classification method are examined both separately and jointly.Conclusions: The median minimal error rate over the investigated classifiers was as low as 31% and 41% based on permuted uninformative predictors from studies on colon cancer and prostate cancer, respectively. We conclude that the strategy to present only the optimal result is not acceptable because it yields a substantial bias in error rate estimation, and suggest alternative approaches for properly reporting classification accuracy.