A multi-loss super regression learner (MSRL) with application to survival prediction using proteomics

A multi-loss super regression learner (MSRL) with application to survival prediction using proteomics
复制标题

DOI:
10.1007/s00180-014-0516-z
复制
发表时间:
2014-12-01
影响因子:
1.3
通讯作者:
Datta, Susmita
Datta, Susmita
中科院分区:
数学4区
文献类型:
--
作者:
Shah, Jasmit;Datta, Somnath;Datta, Susmita

文献摘要

被引文献

相似文献

尽管多年来已经提出了许多回归技术来处理大量的回归变量,但由于最近高通量实验中出现的数据的复杂性,任何单一技术都不可能成功地对所有数据类型进行建模。因此,从现代回归技术的集合,能够处理高维回归的多元回归算法应该被受理分析这样的数据。提出了一种构建超级回归学习器的新方法,该方法可以与训练数据集拟合,以便对连续结果进行未来预测。由此产生的超级回归模型本质上是多目标的,并且无论数据类型如何,都模仿最佳成分回归模型的性能。这是通过结合基于自举的风险计算、等级聚合和堆叠的元素来实现的。这种方法的效用是通过使用质谱数据证明。
Even though a number of regression techniques have been proposed over the years to handle a large number of regressors, due to the complex nature of data emerging from recent high-throughput experiments, it is unlikely that any single technique will be successful in modeling all data types. Thus, multiple regression algorithms from the collection of modern regression techniques that are capable of handling high dimensional regressors should be entertained for analyzing such data. A novel approach of building a super regression learner is proposed which can be fit with a training data set in order to make future predictions of a continuous outcome. The resulting super regression model is multi-objective in nature and mimics the performances of the best component regression models irrespective of the data type. This is accomplished by combining elements of bootstrap based risk calculation, rank aggregation, and stacking. The utility of this approach is demonstrated through its use on mass spectrometry data.