Boosting for high-dimensional two-class prediction.

Boosting for high-dimensional two-class prediction.
复制标题

DOI:
10.1186/s12859-015-0723-9
复制
发表时间:
2015-09-21
期刊:
影响因子:
3
通讯作者:
Lusa L
Lusa L
中科院分区:
生物学4区
文献类型:
--
作者:
Blagus R;Lusa L

文献摘要

被引文献

相似文献

在临床研究中,预测模型用于根据患者的一些特征准确预测患者的结果。对于高维预测模型(变量的数量大大超过样本的数量),选择合适的分类器至关重要,因为据观察,没有一种分类算法能够对所有类型的数据表现最佳。 Boosting 是一种结合使用基分类器获得的分类结果的方法,其中样本权重根据先前迭代的性能顺序调整。一般来说,Boosting 优于任何单个分类器,但对高维数据的研究表明,最标准的 boosting 算法 AdaBoost.M1 无法显着提高其基分类器的性能。最近提出了其他增强算法(梯度增强、随机梯度增强、LogitBoost);它们的性能优于 AdaBoost.M1,但未针对高维数据评估它们的性能。在本文中,我们使用模拟研究和真实的基因表达数据集来评估高维数据时增强算法的性能。我们的结果证实,AdaBoost.M1 在此设置中表现不佳,通常无法提高其基分类器的性能。我们对此进行了解释,并提出了一种修改方案 AdaBoost.M1.ICV,它使用预测误差的交叉验证估计,并且在数据为高维时优于原始算法。当基分类器对训练数据过度拟合时,建议使用 AdaBoost.M1.ICV:变量数量较多、样本数量较少和/或类之间的差异较大。在较小程度上,梯度增强也遇到类似的问题。与低维数据的研究结果相反,当数据是高维时,收缩不会提高梯度提升的性能,但它有利于随机梯度提升,在我们的分析中,随机梯度提升的性能优于其他提升算法。 LogitBoost 存在过度拟合问题,通常表现不佳。结果表明,当数据是高维时,Boosting 也可以显着提高其基分类器的性能。然而,并非所有的 boosting 算法都表现得同样好。 LogitBoost、AdaBoost.M1 和梯度提升对于此类数据似乎不太有用。总体而言,带有收缩的随机梯度提升和 AdaBoost.M1.ICV 似乎是高维类预测的首选。本文的在线版本 (doi:10.1186/s12859-015-0723-9) 包含补充材料,可供授权用户使用。
In clinical research prediction models are used to accurately predict the outcome of the patients based on some of their characteristics. For high-dimensional prediction models (the number of variables greatly exceeds the number of samples) the choice of an appropriate classifier is crucial as it was observed that no single classification algorithm performs optimally for all types of data. Boosting was proposed as a method that combines the classification results obtained using base classifiers, where the sample weights are sequentially adjusted based on the performance in previous iterations. Generally boosting outperforms any individual classifier, but studies with high-dimensional data showed that the most standard boosting algorithm, AdaBoost.M1, cannot significantly improve the performance of its base classier. Recently other boosting algorithms were proposed (Gradient boosting, Stochastic Gradient boosting, LogitBoost); they were shown to perform better than AdaBoost.M1 but their performance was not evaluated for high-dimensional data. In this paper we use simulation studies and real gene-expression data sets to evaluate the performance of boosting algorithms when data are high-dimensional. Our results confirm that AdaBoost.M1 can perform poorly in this setting, often failing to improve the performance of its base classifier. We provide the explanation for this and propose a modification, AdaBoost.M1.ICV, which uses cross-validated estimates of the prediction errors and outperforms the original algorithm when data are high-dimensional. The use of AdaBoost.M1.ICV is advisable when the base classifier overfits the training data: the number of variables is large, the number of samples is small, and/or the difference between the classes is large. To a lesser extent also Gradient boosting suffers from similar problems. Contrary to the findings for the low-dimensional data, shrinkage does not improve the performance of Gradient boosting when data are high-dimensional, however it is beneficial for Stochastic Gradient boosting, which outperformed the other boosting algorithms in our analyses. LogitBoost suffers from overfitting and generally performs poorly. The results show that boosting can substantially improve the performance of its base classifier also when data are high-dimensional. However, not all boosting algorithms perform equally well. LogitBoost, AdaBoost.M1 and Gradient boosting seem less useful for this type of data. Overall, Stochastic Gradient boosting with shrinkage and AdaBoost.M1.ICV seem to be the preferable choices for high-dimensional class-prediction. The online version of this article (doi:10.1186/s12859-015-0723-9) contains supplementary material, which is available to authorized users.