Identifying candidate drivers of drug response in heterogeneous cancer by mining high throughput genomics data.

Identifying candidate drivers of drug response in heterogeneous cancer by mining high throughput genomics data.
复制标题

DOI:
10.1186/s12864-016-2942-5
复制
发表时间:
2016-08-15
期刊:
影响因子:
4.4
通讯作者:
Nabavi S
Nabavi S
中科院分区:
生物学2区
文献类型:
--
作者:
Nabavi S

文献摘要

参考文献

被引文献

相似文献

随着技术的进步,大量的多种类型的高通量基因组学数据可用。这些数据具有巨大的潜力,可以识别新的和有临床价值的生物标志物,以指导诊断,预后评估和复杂疾病(如癌症)的治疗。然而,整合、分析和解释大量嘈杂的基因组数据以获得有生物学意义的结果仍然具有很大的挑战性。利用先进的计算方法挖掘基因组数据集可以帮助解决这些问题。为了便于从大量的异质数据中识别出一系列具有生物学意义的基因作为抗癌药物耐药性的候选驱动因素,我们采用了统计机器学习技术和整合的基因组学数据集。我们开发了一种计算方法,整合了敏感和耐药肿瘤的基因表达,体细胞突变和拷贝数畸变数据。在该方法中,一个基于模块网络分析的综合方法被应用到识别潜在的驱动基因。随后进行交叉验证并比较敏感组和耐药组的结果,以获得候选生物标志物的最终列表。我们将这种方法应用于癌症基因组图谱中的卵巢癌数据。最终结果包含生物学相关基因,如COL 11 A1,在最近的几项研究中,COL 11 A1被报道为上皮性卵巢癌的顺铂耐药生物标志物。所描述的方法产生也控制其共调节基因的表达的异常基因的短列表。结果表明,无偏数据驱动的计算方法可以识别生物相关的候选生物标志物。它可以用于广泛的应用程序,比较两个条件与高度异构的数据集。本文的在线版本(doi:10.1186/s12864-016-2942-5)包含补充材料,可供授权用户使用。
With advances in technologies, huge amounts of multiple types of high-throughput genomics data are available. These data have tremendous potential to identify new and clinically valuable biomarkers to guide the diagnosis, assessment of prognosis, and treatment of complex diseases, such as cancer. Integrating, analyzing, and interpreting big and noisy genomics data to obtain biologically meaningful results, however, remains highly challenging. Mining genomics datasets by utilizing advanced computational methods can help to address these issues. To facilitate the identification of a short list of biologically meaningful genes as candidate drivers of anti-cancer drug resistance from an enormous amount of heterogeneous data, we employed statistical machine-learning techniques and integrated genomics datasets. We developed a computational method that integrates gene expression, somatic mutation, and copy number aberration data of sensitive and resistant tumors. In this method, an integrative method based on module network analysis is applied to identify potential driver genes. This is followed by cross-validation and a comparison of the results of sensitive and resistance groups to obtain the final list of candidate biomarkers. We applied this method to the ovarian cancer data from the cancer genome atlas. The final result contains biologically relevant genes, such as COL11A1, which has been reported as a cis-platinum resistant biomarker for epithelial ovarian carcinoma in several recent studies. The described method yields a short list of aberrant genes that also control the expression of their co-regulated genes. The results suggest that the unbiased data driven computational method can identify biologically relevant candidate biomarkers. It can be utilized in a wide range of applications that compare two conditions with highly heterogeneous datasets. The online version of this article (doi:10.1186/s12864-016-2942-5) contains supplementary material, which is available to authorized users.
DOI: 10.1186/1471-2105-5-110
发表时间: 2004-08-12
期刊: BMC bioinformatics
影响因子: 3
作者:
Lyons-Weiler J;Patel S;Becich MJ;Godfrey TE
通讯作者: Godfrey TE