Integrating single-cell multimodal epigenomic data using 1D-convolutional neural networks.

Integrating single-cell multimodal epigenomic data using 1D-convolutional neural networks.
复制标题

使用一维卷积神经网络整合单细胞多模式表观基因组数据。

DOI:
10.1101/2024.02.16.580655
复制
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Welch,JoshuaD
Welch,JoshuaD
中科院分区:
--
文献类型:
--
作者:
Gao,Chao;Welch,JoshuaD

文献摘要

相似文献

动机最近的实验进展使得单细胞多模式表观基因组分析成为可能,该分析可测量同一细胞内的多个组蛋白修饰和染色质可及性。这种并行测量提供了令人兴奋的新机会来研究表观基因组模式如何在细胞类型和状态之间共同变化。使用这些类型数据的关键步骤是整合表观基因组模式以学习每个细胞的统一表示,但现有方法并非旨在模拟这种数据类型的独特性质。我们的关键见解是将单细胞多模态表观基因组数据建模为多通道序列信号。结果我们开发了ConvNet-VAE,这是一种使用一维(1D)卷积变分自动编码器(VAE)进行单细胞多模态表观基因组数据集成的新颖框架。我们评估了 Nano-CUT&Tag 和单细胞纳米体栓系转座上的 ConvNet-VAE,然后对幼年小鼠大脑和人骨髓生成的测序数据进行了评估。我们发现ConvNet-VAE 可以比以前的架构更好地执行降维和批量校正,同时使用明显更少的参数。此外,卷积和全连接架构之间的性能差距随着模态数量的增加而增加,更深的卷积架构可以提高性能,而更深的全连接架构性能会下降。我们的结果表明,卷积自动编码器是集成当前和未来单细胞多模态表观基因组数据集的一种有前景的方法。可用性和实现VAE 模型的源代码和 Jupyter Notebook 中的演示可在 https://github.com/welch-lab/ConvNetVAE 上获取
MotivationRecent experimental developments enable single-cell multimodal epigenomic profiling, which measures multiple histone modifications and chromatin accessibility within the same cell. Such parallel measurements provide exciting new opportunities to investigate how epigenomic modalities vary together across cell types and states. A pivotal step in using these types of data is integrating the epigenomic modalities to learn a unified representation of each cell, but existing approaches are not designed to model the unique nature of this data type. Our key insight is to model single-cell multimodal epigenome data as a multichannel sequential signal.ResultsWe developedConvNet-VAEs, a novel framework that uses one-dimensional (1D) convolutional variational autoencoders (VAEs) for single-cell multimodal epigenomic data integration. We evaluatedConvNet-VAEs on nano-CUT&Tag and single-cell nanobody-tethered transposition followed by sequencing data generated from juvenile mouse brain and human bone marrow. We found thatConvNet-VAEs can perform dimension reduction and batch correction better than previous architectures while using significantly fewer parameters. Furthermore, the performance gap between convolutional and fully connected architectures increases with the number of modalities, and deeper convolutional architectures can increase the performance, while the performance degrades for deeper fully connected architectures. Our results indicate that convolutional autoencoders are a promising method for integrating current and future single-cell multimodal epigenomic datasets.Availability and implementationThe source code of VAE models and a demo in Jupyter notebook are available at https://github.com/welch-lab/ConvNetVAE