Datamining approaches for modeling tumor control probability

Datamining approaches for modeling tumor control probability
复制标题

DOI:
10.3109/02841861003649224
复制
发表时间:
2010-11-01
期刊:
影响因子:
3.1
通讯作者:
Bradley, Jeffrey D.
Bradley, Jeffrey D.
中科院分区:
医学3区
文献类型:
--
作者:
El Naqa, Issam;Deasy, Joseph O.;Bradley, Jeffrey D.

文献摘要

被引文献

相似文献

背景放疗的肿瘤控制概率(TCP)由肿瘤生物学、肿瘤微环境、辐射剂量学和患者相关变量之间的复杂相互作用决定。这些异质变量相互作用的复杂性构成了为常规临床实践构建预测模型的挑战。我们描述了一个数据挖掘框架,可以解开剂量剂量体积预后变量之间的高阶关系,询问各种放射生物学过程,并推广到以前看不见的数据时,前瞻性地应用。材料和方法。讨论了几种数据挖掘方法,包括剂量-体积度量,等效均匀剂量,机械泊松模型,以及使用统计回归和机器学习技术的模型构建方法。非小细胞肺癌(NSCLC)患者的机构数据集用于证明这些方法。使用双变量斯皮尔曼等级相关(rs)评价不同方法的性能。过拟合是通过抑制方法控制的。结果使用56例原发性NCSLC肿瘤患者和23个候选变量的数据集,我们估计GTV体积和V75是使用统计回归和逻辑模型预测TCP的最佳模型参数。使用这些变量,支持向量机(SVM)内核方法提供了TCP预测的上级性能,与逻辑回归(rs=0.4),基于泊松的TCP(rs=0.33),和细胞杀伤等效均匀剂量模型(rs=0.17)相比,留一法检验的rs=0.68。结论.治疗反应的预测可以通过利用数据挖掘方法来改善,这些方法能够揭示模型变量之间重要的非线性复杂相互作用,并能够预测前瞻性临床应用的未知数据。
Background. Tumor control probability (TCP) to radiotherapy is determined by complex interactions between tumor biology, tumor microenvironment, radiation dosimetry, and patient-related variables. The complexity of these heterogeneous variable interactions constitutes a challenge for building predictive models for routine clinical practice. We describe a datamining framework that can unravel the higher order relationships among dosimetric dose-volume prognostic variables, interrogate various radiobiological processes, and generalize to unseen data before when applied prospectively. Material and methods. Several datamining approaches are discussed that include dose-volume metrics, equivalent uniform dose, mechanistic Poisson model, and model building methods using statistical regression and machine learning techniques. Institutional datasets of non-small cell lung cancer (NSCLC) patients are used to demonstrate these methods. The performance of the different methods was evaluated using bivariate Spearman rank correlations (rs). Over-fitting was controlled via resampling methods. Results. Using a dataset of 56 patients with primary NCSLC tumors and 23 candidate variables, we estimated GTV volume and V75 to be the best model parameters for predicting TCP using statistical resampling and a logistic model. Using these variables, the support vector machine (SVM) kernel method provided superior performance for TCP prediction with an rs=0.68 on leave-one-out testing compared to logistic regression (rs=0.4), Poisson-based TCP (rs=0.33), and cell kill equivalent uniform dose model (rs=0.17). Conclusions. The prediction of treatment response can be improved by utilizing datamining approaches, which are able to unravel important non-linear complex interactions among model variables and have the capacity to predict on unseen data for prospective clinical applications.