CAsubtype: An R Package to Identify Gene Sets Predictive of Cancer Subtypes and Clinical Outcomes

CAsubtype: An R Package to Identify Gene Sets Predictive of Cancer Subtypes and Clinical Outcomes
复制标题

CAsubtype:用于识别预测癌症亚型和临床结果的基因集的 R 包

DOI:
10.1007/s12539-016-0198-z
复制
发表时间:
2018-03-01
影响因子:
4.8
通讯作者:
Li, Hua
Li, Hua
中科院分区:
生物学3区
文献类型:
--
作者:
Kong, Hualei;Tong, Pan;Li, Hua

文献摘要

被引文献

相似文献

在过去的十年中,癌症的分子分类由于与临床实践中常用的传统方法相比其对临床结果的高预测能力而获得了高度普及。特别是,使用基因表达谱,最近的研究已经成功地确定了许多基因集,用于描绘与不同预后相关的癌症亚型。然而,由于缺乏灵活性、集成性和易用性的工具,识别此类基因集仍然是一项艰巨的任务。为了减轻负担,我们开发了一个R包,CAsubtype,以有效地识别预测癌症亚型和临床结果的基因集。通过整合超过13,000个注释的基因集,CAsubtype为新的癌症亚型鉴定提供了一个全面的候选库。为了方便数据访问,CAsubtype还包括来自TCGA的2000多名癌症患者的基因表达和临床数据。CAsubtype首先采用主成分分析来识别基因集(来自用户提供的或包集成的基因集),其中稳健的主成分代表癌症样本之间的显著大的变化。基于这些主成分,CAsubtype在低维空间中可视化样本分布,以便更好地理解样本之间的区别,并使用流行的聚类算法将样本分类为子组。最后,CAsubtype进行生存分析,以比较确定的亚组之间的临床结果,评估其作为潜在的新型癌症亚型的临床价值。总之,CAsubtype是R环境中识别用于癌症亚型识别和临床结果预测的基因集的灵活且良好集成的工具。其简单的R命令和全面的数据集可以有效地检查任何给定基因集的临床价值,从而促进生物学和临床研究中的假设生成和测试。
In the past decade, molecular classification of cancer has gained high popularity owing to its high predictive power on clinical outcomes as compared with traditional methods commonly used in clinical practice. In particular, using gene expression profiles, recent studies have successfully identified a number of gene sets for the delineation of cancer subtypes that are associated with distinct prognosis. However, identification of such gene sets remains a laborious task due to the lack of tools with flexibility, integration and ease of use. To reduce the burden, we have developed an R package, CAsubtype, to efficiently identify gene sets predictive of cancer subtypes and clinical outcomes. By integrating more than 13,000 annotated gene sets, CAsubtype provides a comprehensive repertoire of candidates for new cancer subtype identification. For easy data access, CAsubtype further includes the gene expression and clinical data of more than 2000 cancer patients from TCGA. CAsubtype first employs principal component analysis to identify gene sets (from user-provided or package-integrated ones) with robust principal components representing significantly large variation between cancer samples. Based on these principal components, CAsubtype visualizes the sample distribution in low-dimensional space for better understanding of the distinction between samples and classifies samples into subgroups with prevalent clustering algorithms. Finally, CAsubtype performs survival analysis to compare the clinical outcomes between the identified subgroups, assessing their clinical value as potentially novel cancer subtypes. In conclusion, CAsubtype is a flexible and well-integrated tool in the R environment to identify gene sets for cancer subtype identification and clinical outcome prediction. Its simple R commands and comprehensive data sets enable efficient examination of the clinical value of any given gene set, thus facilitating hypothesis generating and testing in biological and clinical studies.