An empirical assessment of validation practices for molecular classifiers

An empirical assessment of validation practices for molecular classifiers
复制标题

DOI:
10.1093/bib/bbq073
复制
发表时间:
2011-05-01
影响因子:
9.5
通讯作者:
Ioannidis, John P. A.
Ioannidis, John P. A.
中科院分区:
生物学2区
文献类型:
--
作者:
Castaldi, Peter J.;Dahabreh, Issa J.;Ioannidis, John P. A.

文献摘要

被引文献

相似文献

建议的分子分类器可能过于适合噪声基因组和蛋白质组数据的特性。交叉验证方法经常被用来获得分类精度的估计,但模拟和案例研究都表明,当使用不适当的方法时,可能会产生偏差。偏见可以被绕过,泛化能力可以通过外部(独立)验证来测试。我们评估了35项报道了分子分类器的外部验证的研究。我们提取了关于研究设计和方法学特征的信息,并对28项同时进行了内部交叉验证和外部验证的研究比较了分子分类器的性能。我们证明,大多数研究采用的交叉验证做法可能会高估分类器的性能。大多数研究在检测内部交叉验证和外部验证之间的敏感性或特异度下降20%方面明显不足[中位数功率分别为36%(IQR,21-61%)和29%(IQR,15-65%)]。报告的敏感性和特异性的中位数分类性能在交叉验证中分别为94%和98%,在独立验证中分别为88%和81%。交叉验证与独立验证的相对诊断优势比为3.26(95%可信区间2.04~5.21)。最后,我们回顾了所有引用我们研究样本中的那些研究的研究(n=758),并仅确定了这些分类器随后的额外独立验证的一个实例。总之,这些结果证明,文献中采用的许多交叉验证做法可能存在偏见,这一领域的真正进展将需要采用分子分类器的常规外部验证,最好是在比目前实践中大得多的研究中。
Proposed molecular classifiers may be overfit to idiosyncrasies of noisy genomic and proteomic data. Cross-validation methods are often used to obtain estimates of classification accuracy, but both simulations and case studies suggest that, when inappropriate methods are used, bias may ensue. Bias can be bypassed and generalizability can be tested by external (independent) validation. We evaluated 35 studies that have reported on external validation of a molecular classifier. We extracted information on study design and methodological features, and compared the performance of molecular classifiers in internal cross-validation versus external validation for 28 studies where both had been performed. We demonstrate that the majority of studies pursued cross-validation practices that are likely to overestimate classifier performance. Most studies were markedly underpowered to detect a 20% decrease in sensitivity or specificity between internal cross-validation and external validation [median power was 36% (IQR, 21-61%) and 29% (IQR, 15-65%), respectively]. The median reported classification performance for sensitivity and specificity was 94% and 98%, respectively, in cross-validation and 88% and 81% for independent validation. The relative diagnostic odds ratio was 3.26 (95% CI 2.04-5.21) for cross-validation versus independent validation. Finally, we reviewed all studies (n=758) which cited those in our study sample, and identified only one instance of additional subsequent independent validation of these classifiers. In conclusion, these results document that many cross-validation practices employed in the literature are potentially biased and genuine progress in this field will require adoption of routine external validation of molecular classifiers, preferably in much larger studies than in current practice.