Gene expression inference with deep learning

Gene expression inference with deep learning
复制标题

DOI:
10.1093/bioinformatics/btw074
复制
发表时间:
2016-06-15
期刊:
影响因子:
5.8
通讯作者:
Xie, Xiaohui
Xie, Xiaohui
中科院分区:
生物学3区
文献类型:
--
作者:
Chen, Yifei;Li, Yi;Xie, Xiaohui

文献摘要

被引文献

相似文献

动机:大规模的基因表达谱已被广泛用于表征细胞状态,以应对各种疾病条件、遗传扰动等。尽管全基因组表达谱的成本一直在稳步下降,但生成数千个样本的表达谱概要仍然非常昂贵。认识到基因表达往往高度相关,NIH Lincs计划的研究人员开发了一种成本效益高的策略,仅对1000个精心挑选的标志性基因进行图谱分析,并依靠计算方法来推断剩余目标基因的表达。然而,Lincs程序目前采用的计算方法是基于线性回归(LR)的,由于它没有捕捉到基因表达之间复杂的非线性关系,限制了其精度。结果:我们提出了一种深度学习方法(D-Gex)来从标志性基因的表达推断目标基因的表达。我们使用基于微阵列的基因表达综合数据集,包括111K表达谱,来训练我们的模型,并将其性能与其他方法的性能进行比较。在所有基因的平均绝对误差方面,深度学习显著优于LR,相对改善15.33%。基于基因的比较分析表明,深度学习在99.97%的目标基因上的错误率低于LR。我们还在一个独立的基于RNA-Seq的GTEx数据集上测试了我们学习的模型的性能,该数据集由2921个表达谱组成。深度学习仍然以6.57%的相对改善优于LR,并且在81.31%的目标基因上获得了较低的误差。
Motivation: Large-scale gene expression profiling has been widely used to characterize cellular states in response to various disease conditions, genetic perturbations, etc. Although the cost of whole-genome expression profiles has been dropping steadily, generating a compendium of expression profiling over thousands of samples is still very expensive. Recognizing that gene expressions are often highly correlated, researchers from the NIH LINCS program have developed a cost-effective strategy of profiling only similar to 1000 carefully selected landmark genes and relying on computational methods to infer the expression of remaining target genes. However, the computational approach adopted by the LINCS program is currently based on linear regression (LR), limiting its accuracy since it does not capture complex nonlinear relationship between expressions of genes.Results: We present a deep learning method (abbreviated as D-GEX) to infer the expression of target genes from the expression of landmark genes. We used the microarray-based Gene Expression Omnibus dataset, consisting of 111K expression profiles, to train our model and compare its performance to those from other methods. In terms of mean absolute error averaged across all genes, deep learning significantly outperforms LR with 15.33% relative improvement. A gene-wise comparative analysis shows that deep learning achieves lower error than LR in 99.97% of the target genes. We also tested the performance of our learned model on an independent RNA-Seq-based GTEx dataset, which consists of 2921 expression profiles. Deep learning still outperforms LR with 6.57% relative improvement, and achieves lower error in 81.31% of the target genes.