Adversarial deconfounding autoencoder for learning robust gene expression embeddings.

Adversarial deconfounding autoencoder for learning robust gene expression embeddings.
复制标题

DOI:
10.1093/bioinformatics/btaa796
复制
发表时间:
2020-12-30
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Lee, Su-In
Lee, Su-In
中科院分区:
其他
文献类型:
--
作者:
Dincer, Ayse B;Janizek, Joseph D;Lee, Su-In

文献摘要

被引文献

相似文献

动机:基因表达谱数量的增加使得能够使用复杂的模型(例如深度无监督神经网络)从这些谱中提取潜在空间。然而,表达谱,尤其是大量收集时,除了真正感兴趣的信号之外,本身还包含由技术人工制品(例如批次效应)和无趣的生物变量(例如年龄)引入的变异。这些变化源(称为混杂因素)产生的嵌入无法转移到不同的域,即从具有特定混杂因素分布的一个数据集中学习的嵌入不能推广到不同的分布。为了解决这个问题,我们尝试将混杂因素与真实信号分开,以生成生物学信息嵌入。结果:在本文中,我们介绍了对抗性解混自动编码器(AD-AE)方法来解混基因表达潜在空间。 AD-AE 模型由两个神经网络组成:(i) 一个自动编码器,用于生成可以重建原始测量值的嵌入;(ii) 一个经过训练的对手,用于根据该嵌入预测混杂因素。我们联合训练网络来生成嵌入,该嵌入可以编码尽可能多的信息,而无需编码任何混杂信号。通过将 AD-AE 应用于两个不同的基因表达数据集,我们表明我们的模型可以(i)生成不编码混杂信息的嵌入,(ii)保留原始空间中存在的生物信号,以及(iii)成功地跨不同混杂域进行泛化。我们证明 AD-AE 优于标准自动编码器和其他去混杂方法。 可用性和实现:我们的代码和数据可在 https://gitlab.cs.washington.edu/abdincer/ad-ae. 联系方式:补充信息:补充数据可在生物信息学在线获取。
MOTIVATION: Increasing number of gene expression profiles has enabled the use of complex models, such as deep unsupervised neural networks, to extract a latent space from these profiles. However, expression profiles, especially when collected in large numbers, inherently contain variations introduced by technical artifacts (e.g. batch effects) and uninteresting biological variables (e.g. age) in addition to the true signals of interest. These sources of variations, called confounders, produce embeddings that fail to transfer to different domains, i.e. an embedding learned from one dataset with a specific confounder distribution does not generalize to different distributions. To remedy this problem, we attempt to disentangle confounders from true signals to generate biologically informative embeddings.RESULTS: In this article, we introduce the Adversarial Deconfounding AutoEncoder (AD-AE) approach to deconfounding gene expression latent spaces. The AD-AE model consists of two neural networks: (i) an autoencoder to generate an embedding that can reconstruct original measurements, and (ii) an adversary trained to predict the confounder from that embedding. We jointly train the networks to generate embeddings that can encode as much information as possible without encoding any confounding signal. By applying AD-AE to two distinct gene expression datasets, we show that our model can (i) generate embeddings that do not encode confounder information, (ii) conserve the biological signals present in the original space and (iii) generalize successfully across different confounder domains. We demonstrate that AD-AE outperforms standard autoencoder and other deconfounding approaches.AVAILABILITY AND IMPLEMENTATION: Our code and data are available at https://gitlab.cs.washington.edu/abdincer/ad-ae.CONTACT:SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.