Machine learning workflows to estimate class probabilities for precision cancer diagnostics on DNA methylation microarray data

Machine learning workflows to estimate class probabilities for precision cancer diagnostics on DNA methylation microarray data
复制标题

DOI:
10.1038/s41596-019-0251-6
复制
发表时间:
2020-01-13
期刊:
影响因子:
14.8
通讯作者:
Sill, Martin
Sill, Martin
中科院分区:
生物学1区
文献类型:
--
作者:
Maros, Mate E.;Capper, David;Sill, Martin

文献摘要

被引文献

相似文献

基于DNA甲基化数据的精确癌症诊断正在成为分子肿瘤分类的最新技术。选择统计方法的标准,以及校准的概率估计这些典型的高度多类分类任务仍然缺乏。为了支持这一选择,我们评估了成熟的机器学习(ML)分类器,包括随机森林(RF),弹性网络(ELNET),支持向量机(SVM)和提升树与后处理算法的结合,并开发了允许无偏类概率(CP)估计的ML工作流程。校准品包括岭惩罚多项式逻辑回归(MR)和通过拟合逻辑回归(LR)和Firth惩罚LR的Platt标度。我们使用5 x 5倍嵌套交叉验证方案,在最近发表的2,801个样本的脑肿瘤450k DNA甲基化队列中比较了这些工作流程,其中91个诊断类别,并证明了其对癌症基因组图谱外部数据的普遍性。ELNET是具有最佳校准配置文件的顶级独立分类器。最好的整体两阶段工作流程是MR校准的SVM,线性内核紧随其后的是脊校准的调谐RF。对于校准,MR是最有效的,无论主要分类器。由于这些比较而开发的协议为选择ML工作流程及其调整提供了有价值的指导,以使用DNA甲基化数据生成校准良好的CP估计值,用于精确诊断。计算时间因ML算法而异,
DNA methylation data-based precision cancer diagnostics is emerging as the state of the art for molecular tumor classification. Standards for choosing statistical methods with regard to well-calibrated probability estimates for these typically highly multiclass classification tasks are still lacking. To support this choice, we evaluated well-established machine learning (ML) classifiers including random forests (RFs), elastic net (ELNET), support vector machines (SVMs) and boosted trees in combination with post-processing algorithms and developed ML workflows that allow for unbiased class probability (CP) estimation. Calibrators included ridge-penalized multinomial logistic regression (MR) and Platt scaling by fitting logistic regression (LR) and Firth's penalized LR. We compared these workflows on a recently published brain tumor 450k DNA methylation cohort of 2,801 samples with 91 diagnostic categories using a 5 x 5-fold nested cross-validation scheme and demonstrated their generalizability on external data from The Cancer Genome Atlas. ELNET was the top stand-alone classifier with the best calibration profiles. The best overall two-stage workflow was MR-calibrated SVM with linear kernels closely followed by ridge-calibrated tuned RF. For calibration, MR was the most effective regardless of the primary classifier. The protocols developed as a result of these comparisons provide valuable guidance on choosing ML workflows and their tuning to generate well-calibrated CP estimates for precision diagnostics using DNA methylation data. Computation times vary depending on the ML algorithm from