Joint use of over- and under-sampling techniques and cross-validation for the development and assessment of prediction models.

Joint use of over- and under-sampling techniques and cross-validation for the development and assessment of prediction models.
复制标题

DOI:
10.1186/s12859-015-0784-9
复制
发表时间:
2015-11-04
期刊:
影响因子:
3
通讯作者:
Lusa L
Lusa L
中科院分区:
生物学4区
文献类型:
--
作者:
Blagus R;Lusa L

文献摘要

被引文献

相似文献

预测模型用于临床研究,以开发可用于根据患者的某些特征准确预测患者结局的规则。它们在临床医生和卫生政策制定者的决策过程中是一个有价值的工具,因为它们使他们能够估计患者已经或将发展成某种疾病、对治疗有反应或他们的疾病将复发的可能性。在过去的几年里,生物医学界对预测模型的兴趣一直在增长。通常,用于开发预测模型的数据是类别不平衡的,因为只有少数患者经历过该事件(因此属于少数类别)。使用类别不平衡数据开发的预测模型往往在少数类别中实现次优预测精度。这个问题可以通过使用旨在平衡类分布的抽样技术来减少。这些技术包括欠采样和过采样,即在分析中保留多数类样本的一小部分或产生来自少数类的新样本。正确评估预测模型在独立数据上的表现至关重要;在没有独立数据集的情况下,通常使用交叉验证。虽然正确的交叉验证的重要性在生物医学文献中得到了很好的记录,但联合使用采样技术和交叉验证所带来的挑战尚未得到解决。我们表明,必须注意确保对采样数据进行正确的交叉验证,并且当使用过采样技术时,高估预测精度的风险更大。文中给出了基于真实数据集的再分析和仿真研究的实例。我们从生物医学文献中找出了一些结果,其中执行了不正确的交叉验证,我们预计过采样技术的性能被严重高估了。本文的在线版本(doi:10.1186/s12859-0150784-9)包含补充材料,授权用户可以使用。
Prediction models are used in clinical research to develop rules that can be used to accurately predict the outcome of the patients based on some of their characteristics. They represent a valuable tool in the decision making process of clinicians and health policy makers, as they enable them to estimate the probability that patients have or will develop a disease, will respond to a treatment, or that their disease will recur. The interest devoted to prediction models in the biomedical community has been growing in the last few years. Often the data used to develop the prediction models are class-imbalanced as only few patients experience the event (and therefore belong to minority class). Prediction models developed using class-imbalanced data tend to achieve sub-optimal predictive accuracy in the minority class. This problem can be diminished by using sampling techniques aimed at balancing the class distribution. These techniques include under- and oversampling, where a fraction of the majority class samples are retained in the analysis or new samples from the minority class are generated. The correct assessment of how the prediction model is likely to perform on independent data is of crucial importance; in the absence of an independent data set, cross-validation is normally used. While the importance of correct cross-validation is well documented in the biomedical literature, the challenges posed by the joint use of sampling techniques and cross-validation have not been addressed. We show that care must be taken to ensure that cross-validation is performed correctly on sampled data, and that the risk of overestimating the predictive accuracy is greater when oversampling techniques are used. Examples based on the re-analysis of real datasets and simulation studies are provided. We identify some results from the biomedical literature where the incorrect cross-validation was performed, where we expect that the performance of oversampling techniques was heavily overestimated. The online version of this article (doi:10.1186/s12859-015-0784-9) contains supplementary material, which is available to authorized users.