Integrating different data types by regularized unsupervised multiple kernel learning with application to cancer subtype discovery.

Integrating different data types by regularized unsupervised multiple kernel learning with application to cancer subtype discovery.
复制标题

DOI:
10.1093/bioinformatics/btv244
复制
发表时间:
2015-06-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Pfeifer N
Pfeifer N
中科院分区:
其他
文献类型:
--
作者:
Speicher NK;Pfeifer N

文献摘要

被引文献

相似文献

动机:尽管正在进行癌症研究,但可用的治疗方法在数量和有效性方面仍然有限,为个体患者做出治疗决定仍然是一个难题。已建立的子类型有助于指导这些决策,主要基于单个数据类型。然而,涉及各种分子特征测量的多维患者数据的分析可以揭示肿瘤的内在特征。大规模的项目积累了各种癌症类型的此类数据,但我们仍然缺乏计算方法来以有意义的方式可靠地整合这些信息。因此,我们应用和扩展当前的多核学习降维方法。一方面,我们添加了一个正则化项,以避免在优化过程中的过拟合,另一方面,我们表明,甚至可以使用几个内核每个数据类型,从而减轻用户必须选择最好的核函数和核参数为每个数据类型预先。结果:我们已经确定了五种不同癌症类型的生物学意义的亚组。生存分析揭示了所鉴定的亚型的生存时间之间的显著差异,P值与最先进的方法相当甚至更好。此外,我们得到的子类型反映了来自不同数据源的组合模式,并且我们证明了只有很少信息的输入核矩阵对集成核矩阵的影响较小。我们的亚型对特定疗法表现出不同的反应,这最终可能有助于治疗决策。可用性和实现:可执行文件可根据要求提供。联系方式:nora@mpi-inf.mpg.de或npfeifer@mpi-inf.mpg.de
Motivation: Despite ongoing cancer research, available therapies are still limited in quantity and effectiveness, and making treatment decisions for individual patients remains a hard problem. Established subtypes, which help guide these decisions, are mainly based on individual data types. However, the analysis of multidimensional patient data involving the measurements of various molecular features could reveal intrinsic characteristics of the tumor. Large-scale projects accumulate this kind of data for various cancer types, but we still lack the computational methods to reliably integrate this information in a meaningful manner. Therefore, we apply and extend current multiple kernel learning for dimensionality reduction approaches. On the one hand, we add a regularization term to avoid overfitting during the optimization procedure, and on the other hand, we show that one can even use several kernels per data type and thereby alleviate the user from having to choose the best kernel functions and kernel parameters for each data type beforehand. Results: We have identified biologically meaningful subgroups for five different cancer types. Survival analysis has revealed significant differences between the survival times of the identified subtypes, with P values comparable or even better than state-of-the-art methods. Moreover, our resulting subtypes reflect combined patterns from the different data sources, and we demonstrate that input kernel matrices with only little information have less impact on the integrated kernel matrix. Our subtypes show different responses to specific therapies, which could eventually assist in treatment decision making. Availability and implementation: An executable is available upon request. Contact: nora@mpi-inf.mpg.de or npfeifer@mpi-inf.mpg.de