Improved Outcome Prediction Across Data Sources Through Robust Parameter Tuning

Improved Outcome Prediction Across Data Sources Through Robust Parameter Tuning
复制标题

通过稳健的参数调整改进跨数据源的结果预测

DOI:
10.1007/s00357-020-09368-z
复制
发表时间:
2021
影响因子:
2
通讯作者:
R. Hornung
R. Hornung
中科院分区:
计算机科学4区
文献类型:
--
作者:
N. Ellenbach;A.-L. Boulesteix;B. Bischl;K. Unger;R. Hornung

文献摘要

参考文献

被引文献

相似文献

在许多应用领域中,基于高维数据训练的预测规则随后被应用于对来自其他来源的观测进行预测,但它们在这种情况下并不总是表现良好。这是因为来自不同来源的数据集可能具有(略微)不同的分布,即使它们来自相似的人群。在高维数据及更高维度的背景下,大多数预测方法都涉及一个或多个调整参数。它们的值通常通过最大化训练数据的交叉验证预测性能来选择。然而,这个过程隐含地假设最终将应用预测规则的数据遵循与训练数据相同的分布。如果不是这种情况,那么稍微不适合训练数据的较不复杂的预测规则可能是优选的。实际上,调整参数不仅控制预测规则对训练数据的调整程度,而且更一般地控制对训练数据分布的调整程度。基于这一思想,在本文中,我们比较了各种方法,包括新的程序,选择调整参数值,导致更好地推广预测规则比那些基于交叉验证。这些方法中的大多数使用外部验证数据集。在我们基于大量15个转录组数据集的广泛比较研究中,外部数据的调整和具有调整的鲁棒性参数的鲁棒性调整是导致更好地概括预测规则的两种方法。
In many application areas, prediction rules trained based on high-dimensional data are subsequently applied to make predictions for observations from other sources, but they do not always perform well in this setting. This is because data sets from different sources can feature (slightly) differing distributions, even if they come from similar populations. In the context of high-dimensional data and beyond, most prediction methods involve one or several tuning parameters. Their values are commonly chosen by maximizing the cross-validated prediction performance on the training data. This procedure, however, implicitly presumes that the data to which the prediction rule will be ultimately applied, follow the same distribution as the training data. If this is not the case, less complex prediction rules that slightly underfit the training data may be preferable. Indeed, a tuning parameter does not only control the degree of adjustment of a prediction rule to the training data, but also, more generally, the degree of adjustment to thedistribution ofthe training data. On the basis of this idea, in this paper we compare various approaches including new procedures for choosing tuning parameter values that lead to better generalizing prediction rules than those obtained based on cross-validation. Most of these approaches use an external validation data set. In our extensive comparison study based on a large collection of 15 transcriptomic data sets, tuning on external data and robust tuning with a tuned robustness parameter are the two approaches leading to better generalizing prediction rules.
DOI: --
发表时间: --
期刊:
影响因子: --
作者:
S. Gottlieb;D. Gottlieb;Chi
通讯作者: Chi
DOI: --
发表时间: 2000
期刊:
影响因子: --
作者:
Vladimir;VapnikAT
通讯作者: VapnikAT
DOI: 10.1093/bioinformatics/btu279
发表时间: 2014-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Bernau C;Riester M;Boulesteix AL;Parmigiani G;Huttenhower C;Waldron L;Trippa L
通讯作者: Trippa L
微阵列实验中的批次效应和噪声
DOI: --
发表时间: 2009
期刊:
影响因子: --
作者:
A. Scherer
通讯作者: A. Scherer
DOI: 10.18637/jss.v033.i01
发表时间: 2010-02-01
影响因子: 5.8
作者:
Friedman, Jerome;Hastie, Trevor;Tibshirani, Rob
通讯作者: Tibshirani, Rob