A kernel-based integration of genome-wide data for clinical decision support

A kernel-based integration of genome-wide data for clinical decision support
复制标题

DOI:
10.1186/gm39
复制
发表时间:
2009-01-01
期刊:
影响因子:
12.3
通讯作者:
De Moor, Bart
De Moor, Bart
中科院分区:
生物学1区
文献类型:
--
作者:
Daemen, Anneleen;Gevaert, Olivier;De Moor, Bart

文献摘要

被引文献

相似文献

背景资料:尽管微阵列技术允许在一个实验中研究肿瘤的转录组组成,但由于可变剪接、翻译后修饰以及病理条件(例如癌症)对转录和翻译的影响,转录组不能完全反映潜在的生物学。这增加了融合多个全基因组数据源的重要性,例如基因组、转录组、蛋白质组和表观基因组。目前增加的可用组学数据的量强调需要一个方法集成framework.Methods:我们提出了一个基于内核的方法,许多全基因组数据源相结合的临床决策支持。在构建分类器之前,在核心矩阵的水平上在患者域内进行整合。作为监督分类算法,使用加权最小二乘支持向量机。我们将此框架应用于两个癌症病例,即,包含微阵列和蛋白质组学数据的直肠癌数据集和包含微阵列和基因组学数据的前列腺癌数据集。对于这两种情况下,多个outcomes是predicted.Results:对于直肠癌的结果,最高的留一法(LOO)下的受试者工作特征曲线(AUC)的面积时,结合治疗过程中收集的微阵列和蛋白质组学数据,范围从0.927到0.987。对于前列腺癌,所有四个结果有一个更好的LOO AUC相结合时,微阵列和基因组学数据,范围从0.786复发到0.987的metastasis.Conclusions:对于这两个癌症网站的预测,所有的结果改善时,一个以上的全基因组数据集被认为是。这表明,整合多个全基因组数据源可以提高临床决策支持模型的预测性能。这就强调需要全面的多模式数据。我们承认,在第一阶段,这将大大增加成本;然而,这是一项必要的投资,以最终获得可用于患者定制治疗的具有成本效益的模型。
Background: Although microarray technology allows the investigation of the transcriptomic make-up of a tumor in one experiment, the transcriptome does not completely reflect the underlying biology due to alternative splicing, post-translational modifications, as well as the influence of pathological conditions (for example, cancer) on transcription and translation. This increases the importance of fusing more than one source of genome-wide data, such as the genome, transcriptome, proteome, and epigenome. The current increase in the amount of available omics data emphasizes the need for a methodological integration framework.Methods: We propose a kernel-based approach for clinical decision support in which many genome-wide data sources are combined. Integration occurs within the patient domain at the level of kernel matrices before building the classifier. As supervised classification algorithm, a weighted least squares support vector machine is used. We apply this framework to two cancer cases, namely, a rectal cancer data set containing microarray and proteomics data and a prostate cancer data set containing microarray and genomics data. For both cases, multiple outcomes are predicted.Results: For the rectal cancer outcomes, the highest leave-one-out (LOO) areas under the receiver operating characteristic curves (AUC) were obtained when combining microarray and proteomics data gathered during therapy and ranged from 0.927 to 0.987. For prostate cancer, all four outcomes had a better LOO AUC when combining microarray and genomics data, ranging from 0.786 for recurrence to 0.987 for metastasis.Conclusions: For both cancer sites the prediction of all outcomes improved when more than one genome-wide data set was considered. This suggests that integrating multiple genome-wide data sources increases the predictive performance of clinical decision support models. This emphasizes the need for comprehensive multi-modal data. We acknowledge that, in a first phase, this will substantially increase costs; however, this is a necessary investment to ultimately obtain cost-efficient models usable in patient tailored therapy.