The C1C2: a framework for simultaneous model selection and assessment.

The C1C2: a framework for simultaneous model selection and assessment.
复制标题

DOI:
10.1186/1471-2105-9-360
复制
发表时间:
2008-09-02
期刊:
影响因子:
3
通讯作者:
Wikberg, Jarl E. S.
Wikberg, Jarl E. S.
中科院分区:
生物学4区
文献类型:
--
作者:
Eklund, Martin;Spjuth, Ola;Wikberg, Jarl E. S.

文献摘要

参考文献

被引文献

相似文献

最近,人们一直担心预测建模方法无法推广到新数据。其中一些问题可以归因于模型选择和评估方法不当。在这里,我们通过引入一个新颖的通用框架C1C2来解决这个问题,该框架用于同时选择和评估模型。该框架依赖于对数据的划分,以便根据使用的数据将模型选择与模型评估分开。由于通常可以想到的模型的数量是巨大的,因此研究两种自动搜索方法--遗传算法和蛮力方法--用于模型选择也是很有意义的。作为演示,C1C2被应用于模拟和真实世界的数据集。假设惩罚线性模型合理地逼近因变量和自变量之间的真实关系,从而将模型选择问题归结为变量选择和惩罚参数的选择问题。我们还研究了假设相关变量个数的先验知识对模型选择和泛化误差估计的影响。将使用C1C2获得的结果与采用重复K-折叠交叉验证来选择和评估模型所获得的结果进行比较。C1C2框架在选择正确的变量子集和为惩罚参数产生合理选择方面做得很好,即使在自变量高度相关和观察数量少于变量数量的情况下也是如此。C1C2框架也被发现给出了对泛化误差的准确估计。关于重要自变量个数的先验信息改善了变量子集的选择,但降低了泛化误差估计的精度。与使用强力方法相比,使用遗传算法使模型选择变差,但不会使泛化误差估计变差。在模型选择方面,重复K-折叠交叉验证得到的结果与C1C2得到的结果相似,但普遍误差估计的准确性较低。C1C2框架被证明可以很好地在惩罚线性模型类中找到真正的模型,并准确地评估其泛化误差,即使对于具有许多高度相关的自变量、低观测变量比和模型假设偏差的数据集也是如此。根据用于每项任务的数据将模型选择和模型评估完全分开,改进了对泛化误差的估计。
There has been recent concern regarding the inability of predictive modeling approaches to generalize to new data. Some of the problems can be attributed to improper methods for model selection and assessment. Here, we have addressed this issue by introducing a novel and general framework, the C1C2, for simultaneous model selection and assessment. The framework relies on a partitioning of the data in order to separate model choice from model assessment in terms of used data. Since the number of conceivable models in general is vast, it was also of interest to investigate the employment of two automatic search methods, a genetic algorithm and a brute-force method, for model choice. As a demonstration, the C1C2 was applied to simulated and real-world datasets. A penalized linear model was assumed to reasonably approximate the true relation between the dependent and independent variables, thus reducing the model choice problem to a matter of variable selection and choice of penalizing parameter. We also studied the impact of assuming prior knowledge about the number of relevant variables on model choice and generalization error estimates. The results obtained with the C1C2 were compared to those obtained by employing repeated K-fold cross-validation for choosing and assessing a model. The C1C2 framework performed well at finding the true model in terms of choosing the correct variable subset and producing reasonable choices for the penalizing parameter, even in situations when the independent variables were highly correlated and when the number of observations was less than the number of variables. The C1C2 framework was also found to give accurate estimates of the generalization error. Prior information about the number of important independent variables improved the variable subset choice but reduced the accuracy of generalization error estimates. Using the genetic algorithm worsened the model choice but not the generalization error estimates, compared to using the brute-force method. The results obtained with repeated K-fold cross-validation were similar to those produced by the C1C2 in terms of model choice, however a lower accuracy of the generalization error estimates was observed. The C1C2 framework was demonstrated to work well for finding the true model within a penalized linear model class and accurately assess its generalization error, even for datasets with many highly correlated independent variables, a low observation-to-variable ratio, and model assumption deviations. A complete separation of the model choice and the model assessment in terms of data used for each task improves the estimates of the generalization error.
DOI: 10.1021/jm00163a023
发表时间: 1990-01-01
影响因子: 7.3
作者:
SELWOOD, DL;LIVINGSTONE, DJ;STABLES, JN
通讯作者: STABLES, JN
DOI: 10.1016/s0140-6736(05)17866-0
发表时间: 2005-02-05
期刊: LANCET
影响因子: 168.9
作者:
Michiels, S;Koscielny, S;Hill, C
通讯作者: Hill, C
DOI: 10.1080/00401706.1970.10488634
发表时间: 1970-01-01
期刊: TECHNOMETRICS
影响因子: 2.5
作者:
HOERL, AE;KENNARD, RW
通讯作者: KENNARD, RW
DOI: 10.1186/1471-2105-6-241
发表时间: 2005-10-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Freyhult E;Gardner PP;Moulton V
通讯作者: Moulton V
DOI: 10.1109/tac.1974.1100705
发表时间: 1974-01-01
影响因子: 6.8
作者:
AKAIKE, H
通讯作者: AKAIKE, H