Overall Survival Prognostic Modelling of Non-small Cell Lung Cancer Patients Using Positron Emission Tomography/Computed Tomography Harmonised Radiomics Features: The Quest for the Optimal Machine Learning Algorithm

Overall Survival Prognostic Modelling of Non-small Cell Lung Cancer Patients Using Positron Emission Tomography/Computed Tomography Harmonised Radiomics Features: The Quest for the Optimal Machine Learning Algorithm
复制标题

DOI:
10.1016/j.clon.2021.11.014
复制
发表时间:
2022-01-17
期刊:
影响因子:
3.4
通讯作者:
Zaidi, Habib
Zaidi, Habib
中科院分区:
医学2区
文献类型:
--
作者:
Amini, Mehdi;Hajianfar, Ghasem;Zaidi, Habib

文献摘要

被引文献

相似文献

目的:尽管放射组学预后模型在各种临床应用中取得了令人鼓舞的结果,但仍需要解决多个挑战。放射组学预后模型的两个主要局限性包括由于单一成像方式导致的信息限制以及针对所考虑的方式和临床结果选择最佳机器学习和特征选择方法。在这项工作中,我们将几种特征选择和机器学习方法应用于单模态正电子发射断层扫描(PET)和计算机断层扫描(CT)以及多模态PET/CT融合,以确定不同放射组学模态的最佳组合,从而预测非小细胞肺癌患者的总体生存率。材料和方法:本研究使用了来自癌症成像档案的PET/CT数据集,包括来自两个独立机构的受试者(87名和95名患者)。每个队列使用一次作为训练,一次作为测试,然后对结果取平均值。ComBat协调用于解决中心效应。在我们提出的放射组学框架中,除了单模态PET和CT模型外,多模态放射组学模型使用多级(特征和图像级别)融合开发。针对特征级策略考虑了两种不同的方法,包括将PET和CT特征连接到单个特征集中并交替地对它们进行平均。对于图像级融合,我们使用了三种不同的融合方法,即小波融合,基于引导滤波的融合和潜在的低秩表示融合。在所提出的预后建模框架中,将四种特征选择和七种机器学习方法的组合应用于所有放射组学模式(两种单一模式和五种多模式),优化机器学习超参数,最后通过自举在测试队列中进行1000次重复评估模型。特征选择和机器学习方法被选为文献中的流行技术,由公共领域的开源软件及其科普连续事件生存数据的能力提供支持。采用多因素方差分析进行变异性分析,并通过偏差校正效应量估计值u(2)计算由放射组学模式、特征选择和机器学习方法解释的总方差的比例。然而,最小深度(MD)作为特征选择和Lasso和Elastic-Net正则化广义线性模型(glmnet)作为机器学习方法具有最高的平均结果。ANOVA检验的结果表明,每个因素(放射组学模态、特征选择和机器学习方法)引入模型性能的变异性是特定于情况的,即,不同放射组学模态和融合策略的方差不同。总体而言,最大比例的方差解释了机器学习,除了在功能级别的fusion strategy.Conclusion模型:最佳特征选择和机器学习方法的识别是一个关键的一步,在发展健全和准确的放射组学风险模型。此外,最佳方法是病例特异性的,由于所使用的放射组学模态和融合策略而不同。(C)2021作者由Elsevier Ltd代表皇家放射科医师学院发布。
Aims: Despite the promising results achieved by radiomics prognostic models for various clinical applications, multiple challenges still need to be addressed. The two main limitations of radiomics prognostic models include information limitation owing to single imaging modalities and the selection of optimum machine learning and feature selection methods for the considered modality and clinical outcome. In this work, we applied several feature selection and machine learning methods to single-modality positron emission tomography (PET) and computed tomography (CT) and multimodality PET/CT fusion to identify the best combinations for different radiomics modalities towards overall survival prediction in non-small cell lung cancer patients.Materials and methods: A PET/CT dataset from The Cancer Imaging Archive, including subjects from two independent institutions (87 and 95 patients), was used in this study. Each cohort was used once as training and once as a test, followed by averaging of the results. ComBat harmonisation was used to address the centre effect. In our proposed radiomics framework, apart from single-modality PET and CT models, multimodality radiomics models were developed using multilevel (feature and image levels) fusion. Two different methods were considered for the feature-level strategy, including concatenating PET and CT features into a single feature set and alternatively averaging them. For image-level fusion, we used three different fusion methods, namely wavelet fusion, guided filtering-based fusion and latent low-rank representation fusion. In the proposed prognostic modelling framework, combinations of four feature selection and seven machine learning methods were applied to all radiomics modalities (two single and five multimodalities), machine learning hyper-parameters were optimised and finally the models were evaluated in the test cohort with 1000 repetitions via bootstrapping. Feature selection and machine learning methods were selected as popular techniques in the literature, supported by open source software in the public domain and their ability to cope with continuous time-to-event survival data. Multifactor ANOVA was used to carry out variability analysis and the proportion of total variance explained by radiomics modality, feature selection and machine learning methods was calculated by a bias-corrected effect size estimate known as u(2).Results: Optimum feature selection and machine learning methods differed owing to the applied radiomics modality. However, minimum depth (MD) as feature selection and Lasso and Elastic-Net regularized generalized linear model (glmnet) as machine learning method had the highest average results. Results from the ANOVA test indicated that the variability that each factor (radiomics modality, feature selection and machine learning methods) introduces to the performance of models is case specific, i.e. variances differ regarding different radiomics modalities and fusion strategies. Overall, the greatest proportion of variance was explained by machine learning, except for models in feature-level fusion strategy.Conclusion: The identification of optimal feature selection and machine learning methods is a crucial step in developing sound and accurate radiomics risk models. Furthermore, optimum methods are case specific, differing due to the radiomics modality and fusion strategy used. (C) 2021 The Authors. Published by Elsevier Ltd on behalf of The Royal College of Radiologists.