Sparse kernel methods for high-dimensional survival data

Sparse kernel methods for high-dimensional survival data
复制标题

DOI:
10.1093/bioinformatics/btn253
复制
发表时间:
2008-07-15
期刊:
影响因子:
5.8
通讯作者:
Messow, Claudia-Martina
Messow, Claudia-Martina
中科院分区:
生物学3区
文献类型:
--
作者:
Evers, Ludger;Messow, Claudia-Martina

文献摘要

被引文献

相似文献

支持向量机(SVM)等稀疏核方法已经成功地应用于分类和(标准)回归设置。然而,现有的支持向量分类和回归技术并不适合部分删减的生存数据,这些数据通常使用cox比例风险模型进行分析。由于比例风险模型的部分似然仅通过内积依赖于协变量,因此可以进行核化。然而,核比例风险模型产生了一个密集的解决方案,即解决方案取决于所有的观测值。支持向量机的一个关键特征是,它产生一个稀疏的解决方案,只依赖于一小部分训练数据。我们提出两种方法。一种是基于几何思想,类似于支持向量分类,失败的观察和当前有风险的观察之间的余量被最大化。另一种方法是通过一个接一个地添加观测值来获得稀疏模型,类似于导入向量机(IVM)。研究的数据示例表明,这两种方法都可以优于竞争方法。
Sparse kernel methods like support vector machines (SVM) have been applied with great success to classification and (standard) regression settings. Existing support vector classification and regression techniques however are not suitable for partly censored survival data, which are typically analysed using Coxs proportional hazards model. As the partial likelihood of the proportional hazards model only depends on the covariates through inner products, it can be kernelized. The kernelized proportional hazards model however yields a solution that is dense, i.e. the solution depends on all observations. One of the key features of an SVM is that it yields a sparse solution, depending only on a small fraction of the training data. We propose two methods. One is based on a geometric idea, whereakin to support vector classificationthe margin between the failed observation and the observations currently at risk is maximised. The other approach is based on obtaining a sparse model by adding observations one after another akin to the Import Vector Machine (IVM). Data examples studied suggest that both methods can outperform competing approaches.