Transforming L1000 profiles to RNA-seq-like profiles with deep learning.

Transforming L1000 profiles to RNA-seq-like profiles with deep learning.
复制标题

DOI:
10.1186/s12859-022-04895-5
复制
发表时间:
2022-09-13
期刊:
影响因子:
3
通讯作者:
--
中科院分区:
生物学4区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

L1000技术是一种具有成本效益的高通量转录组学技术,已被应用于分析人类细胞系的基因表达对bbb30 000化学和遗传扰动的反应。目前总共有超过300万个可用的L1000配置文件。这样的数据集对于发现药物和候选靶标以及推断小分子的作用机制是非常宝贵的。L1000检测仅测量978个标志性基因的mRNA表达,而计算可靠地推断了11,350个其他基因。缺乏全基因组覆盖限制了对一半人类蛋白质编码基因的知识发现,以及与其他转录组学分析数据整合的潜力。在这里,我们提出了一个深度学习两步模型,将L1000谱转换为rna -seq样谱。模型的输入是测量到的978个标志性基因,而输出是23,614个rna -seq样基因表达谱的载体。该模型首先使用应用于未配对数据的改进CycleGAN模型将具有里程碑意义的基因转化为rna -seq样的978基因谱。然后用全连接的神经网络模型将转化的978个rna -seq样地标基因外推到全基因组空间中。在LINCS和GTEx程序生成的已发表的配对L1000/RNA-seq数据集上,两步模型的Pearson相关系数为0.914,均方根误差为1.167。经过处理的rna -seq样谱可供下载,签名搜索和以独特案例研究为中心的基因反向搜索。在线版本包含补充材料,可在10.1186/s12859-022-04895-5获得。
The L1000 technology, a cost-effective high-throughput transcriptomics technology, has been applied to profile a collection of human cell lines for their gene expression response to > 30,000 chemical and genetic perturbations. In total, there are currently over 3 million available L1000 profiles. Such a dataset is invaluable for the discovery of drug and target candidates and for inferring mechanisms of action for small molecules. The L1000 assay only measures the mRNA expression of 978 landmark genes while 11,350 additional genes are computationally reliably inferred. The lack of full genome coverage limits knowledge discovery for half of the human protein coding genes, and the potential for integration with other transcriptomics profiling data. Here we present a Deep Learning two-step model that transforms L1000 profiles to RNA-seq-like profiles. The input to the model are the measured 978 landmark genes while the output is a vector of 23,614 RNA-seq-like gene expression profiles. The model first transforms the landmark genes into RNA-seq-like 978 gene profiles using a modified CycleGAN model applied to unpaired data. The transformed 978 RNA-seq-like landmark genes are then extrapolated into the full genome space with a fully connected neural network model. The two-step model achieves 0.914 Pearson’s correlation coefficients and 1.167 root mean square errors when tested on a published paired L1000/RNA-seq dataset produced by the LINCS and GTEx programs. The processed RNA-seq-like profiles are made available for download, signature search, and gene centric reverse search with unique case studies. The online version contains supplementary material available at 10.1186/s12859-022-04895-5.
DOI: 10.1093/nar/gkw377
发表时间: 2016-07-08
影响因子: 14.9
作者:
Kuleshov MV;Jones MR;Rouillard AD;Fernandez NF;Duan Q;Wang Z;Koplev S;Jenkins SL;Jagodnik KM;Lachmann A;McDermott MG;Monteiro CD;Gundersen GW;Ma'ayan A
通讯作者: Ma'ayan A
DOI: 10.1186/1471-2105-15-79
发表时间: 2014-03-21
期刊: BMC bioinformatics
影响因子: 3
作者:
Clark NR;Hu KS;Feldmann AS;Kou Y;Chen EY;Duan Q;Ma'ayan A
通讯作者: Ma'ayan A
DOI: 10.1146/annurev-pathol-121808-102144
发表时间: 2010
期刊: Annual review of pathology
影响因子: --
作者:
Coppé JP;Desprez PY;Krtolica A;Campisi J
通讯作者: Campisi J
DOI: 10.1093/bioinformatics/btq466
发表时间: 2010-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Lachmann, Alexander;Xu, Huilei;Ma'ayan, Avi
通讯作者: Ma'ayan, Avi
DOI: 10.1158/1078-0432.ccr-20-0446
发表时间: 2020-11-01
期刊: Clinical cancer research : an official journal of the American Association for Cancer Research
影响因子: --
作者:
Fane ME;Ecker BL;Kaur A;Marino GE;Alicea GM;Douglass SM;Chhabra Y;Webster MR;Marshall A;Colling R;Espinosa O;Coupe N;Maroo N;Campo L;Middleton MR;Corrie P;Xu X;Karakousis GC;Weeraratna AT
通讯作者: Weeraratna AT