Comprehensive evaluation of composite gene features in cancer outcome prediction.

Comprehensive evaluation of composite gene features in cancer outcome prediction.
复制标题

DOI:
10.4137/cin.s14028
复制
发表时间:
2014
期刊:
影响因子:
2
通讯作者:
Koyutürk M
Koyutürk M
中科院分区:
其他
文献类型:
--
作者:
Hou D;Koyutürk M

文献摘要

被引文献

相似文献

由于癌症的异质性和不断演变的性质,基于单个基因表达的分类器通常不会导致癌症结果的稳健预测。作为替代方案,已经提出了联合收割机组合功能相关基因的复合基因特征。预期这样的特征可以是更稳健和可再现的,因为它们可以捕获相关生物过程中作为整体的改变,并且可能对单个基因表达的波动不太敏感。人们已经开发了各种算法来识别复合特征和推断复合基因特征活性,这些算法都声称可以提高预测准确性。然而,由于每个单独的研究所包含的测试数据集的限制和不一致的测试程序,这些研究的结果有时是相互矛盾的和不可生产的。因此,很难全面了解复合基因特征的预测性能,特别是在不同的癌症、癌症亚型和队列中。在这项研究中,我们实现了各种算法的识别复合基因的功能和它们在癌症预后预测的利用,并进行广泛的比较和评估,使用七个微阵列数据集,涵盖两种癌症类型和三种不同的表型。我们的研究结果表明,虽然某些算法优于其他算法的某些分类任务,没有一个算法始终优于其他算法和个别基因功能。
Owing to the heterogeneous and continuously evolving nature of cancers, classifiers based on the expression of individual genes usually do not result in robust prediction of cancer outcome. As an alternative, composite gene features that combine functionally related genes have been proposed. It is expected that such features can be more robust and reproducible since they can capture the alterations in relevant biological processes as a whole and may be less sensitive to fluctuations in the expression of individual genes. Various algorithms have been developed for the identification of composite features and inference of composite gene feature activity, which all claim to improve the prediction accuracy. However, because of the limitations of test datasets incorporated by each individual study and inconsistent test procedures, the results of these studies are sometimes conflicting and unproducible. For this reason, it is difficult to have a comprehensive understanding of the prediction performance of composite gene features, particularly across different cancers, cancer subtypes, and cohorts. In this study, we implement various algorithms for the identification of composite gene features and their utilization in cancer outcome prediction, and perform extensive comparison and evaluation using seven microarray datasets covering two cancer types and three different phenotypes. Our results show that, while some algorithms outperform others for certain classification tasks, no single algorithm consistently outperforms other algorithms and individual gene features.