Determining relative importance of variables in developing and validating predictive models

Determining relative importance of variables in developing and validating predictive models
复制标题

DOI:
10.1186/1471-2288-9-64
复制
发表时间:
2009-09-14
影响因子:
4
通讯作者:
Sung, Lillian
Sung, Lillian
中科院分区:
医学3区
文献类型:
--
作者:
Beyene, Joseph;Atenafu, Eshetu G.;Sung, Lillian

文献摘要

被引文献

相似文献

背景:多元回归模型在广泛的科学学科中使用,自动模型选择程序经常被用来识别独立的预测因素。然而,潜在预测因素的相对重要性的确定以及对其稳定性、预测准确性和泛化能力的验证往往被忽视或没有进行彻底的验证。方法:以一个旨在预测患肿瘤溶解综合征(TLS)风险较低的急性淋巴细胞性白血病(ALL)儿童为研究对象的案例研究,我们提出并比较了两种策略,即自举和随机数据分割,根据潜在预测因素相对于模型稳定性和泛化能力的相对重要性对其进行排序。我们还提出了一种基于解释变异百分比和受试者工作特征(ROC)曲线下面积的相对增加的方法来建立模型,其中排序列表中的变量根据它们的重要性进入模型。另一个旨在确定前列腺癌渗透的预测因素的数据集也用于说明目的。结果:年龄被选为TLS最重要的预测因素。它是使用自举方法100%选择的。使用随机分裂方法,在训练数据中99%的时间被选择,并且在验证数据集中98%的时间是显著的(在5%水平上)。这表明年龄是TLS的一个稳定的预测因子,具有较好的泛化能力。第二个最重要的变量是白细胞计数(WBC)。我们的方法还确定了TLS的一个重要预测因子,如果仅依赖任何自动模型选择程序,该因子将被忽略。低风险组包括10岁以下,无T细胞免疫表型的儿童,其基线白细胞为<20×10(9)/L,可触及的脾为<2 cm。对于前列腺癌数据集,Gleason评分和直肠指诊被认为是肿瘤是否穿透前列腺囊的最重要的指标。结论:我们基于Bootstrap重新采样和重复随机分割技术的模型选择程序可以用来评估证据的强度,即变量确实是一个独立的、可重复性的预测因子。因此,我们的方法可以用于开发性能良好的稳定的、可重复性的模型。此外,我们的方法可以作为一个很好的工具来验证预测模型。之前的生物学和临床研究支持基于我们的选择和验证策略的结果。然而,可能需要大量的模拟来评估我们的方法在不同场景下的性能,以及检查它们对数据中随机波动的敏感度。
Background: Multiple regression models are used in a wide range of scientific disciplines and automated model selection procedures are frequently used to identify independent predictors. However, determination of relative importance of potential predictors and validating the fitted models for their stability, predictive accuracy and generalizability are often overlooked or not done thoroughly.Methods: Using a case study aimed at predicting children with acute lymphoblastic leukemia (ALL) who are at low risk of Tumor Lysis Syndrome (TLS), we propose and compare two strategies, bootstrapping and random split of data, for ordering potential predictors according to their relative importance with respect to model stability and generalizability. We also propose an approach based on relative increase in percentage of explained variation and area under the Receiver Operating Characteristic (ROC) curve for developing models where variables from our ordered list enter the model according to their importance. An additional data set aimed at identifying predictors of prostate cancer penetration is also used for illustrative purposes.Results: Age is chosen to be the most important predictor of TLS. It is selected 100% of the time using the bootstrapping approach. Using the random split method, it is selected 99% of the time in the training data and is significant (at 5% level) 98% of the time in the validation data set. This indicates that age is a stable predictor of TLS with good generalizability. The second most important variable is white blood cell count (WBC). Our methods also identified an important predictor of TLS that was otherwise omitted if relying on any of the automated model selection procedures alone. A group at low risk of TLS consists of children younger than 10 years of age, without T-cell immunophenotype, whose baseline WBC is < 20 x 10(9)/L and palpable spleen is < 2 cm. For the prostate cancer data set, the Gleason score and digital rectal exam are identified to be the most important indicators of whether tumor has penetrated the prostate capsule.Conclusion: Our model selection procedures based on bootstrap re-sampling and repeated random split techniques can be used to assess the strength of evidence that a variable is truly an independent and reproducible predictor. Our methods, therefore, can be used for developing stable and reproducible models with good performances. Moreover, our methods can serve as a good tool for validating a predictive model. Previous biological and clinical studies support the findings based on our selection and validation strategies. However, extensive simulations may be required to assess the performance of our methods under different scenarios as well as check their sensitivity to a random fluctuation in the data.