Multimodal Single-Cell Translation and Alignment with Semi-Supervised Learning.

Multimodal Single-Cell Translation and Alignment with Semi-Supervised Learning.
复制标题

多模式单细胞翻译和半监督学习比对。

DOI:
10.1089/cmb.2022.0264
复制
发表时间:
2022
期刊:
Journal of computational biology : a journal of computational molecular cell biology
影响因子:
--
通讯作者:
Noble,WilliamStafford
Noble,WilliamStafford
中科院分区:
--
文献类型:
--
作者:
Zhang,Ran;Meng-Papaxanthos,Laetitia;Vert,Jean-Philippe;Noble,WilliamStafford

文献摘要

相似文献

单细胞多组学技术能够对细胞调控进行全面的研究,然而大多数单细胞分析只能测量每个细胞的一种活性,如转录、染色质可及性、DNA甲基化或三维染色质结构。为了实现单个细胞的多模态视图,我们提出了Polarbear,这是一种半监督机器学习框架,可以促进缺失模态轮廓预测和单细胞跨模态对齐。Polarbear通过使用来自联合测定的数据以及公共数据库中大量可用的单次测定数据来学习在模式之间进行转换。这种半监督方案减轻了与低细胞数量和高稀疏性相关的问题。Polarbear首先使用联合分析和单分析配置文件为每种模式预训练β变分自编码器,以学习单个细胞的鲁棒表示,然后使用联合分析标签来训练这些细胞表示之间的翻译器。与完全监督的方法相比,这种半监督框架使我们能够预测缺失的模态剖面,并以更高的精度匹配跨模态的单个细胞,从而促进多模态数据集成。
Single-cell multi-omics technologies enable comprehensive interrogation of cellular regulation, yet most single-cell assays measure only one type of activity—such as transcription, chromatin accessibility, DNA methylation, or 3D chromatin architecture—for each cell. To enable a multimodal view for individual cells, we propose Polarbear, a semi-supervised machine learning framework that facilitates missing modality profile prediction and single-cell cross-modality alignment. Polarbear learns to translate between modalities by using data from co-assay measurements coupled with the large quantity of single-assay data available in public databases. This semi-supervised scheme mitigates issues related to low cell quantities and high sparsity in co-assay data. Polarbear first pre-trains a beta-variational autoencoder for each modality using both co-assay and single-assay profiles to learn robust representations of individual cells, and it then uses the co-assay labels to train a translator between these cell representations. This semi-supervised framework enables us to predict missing modality profiles and match single cells across modalities with improved accuracy compared with fully supervised methods, thus facilitating multimodal data integration.