Celda: a Bayesian model to perform co-clustering of genes into modules and cells into subpopulations using single-cell RNA-seq data.

Celda: a Bayesian model to perform co-clustering of genes into modules and cells into subpopulations using single-cell RNA-seq data.
复制标题

DOI:
10.1093/nargab/lqac066
复制
发表时间:
2022-09
影响因子:
4.6
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

单细胞RNA-seq(scRNA-seq)已经成为一种强大的技术,可以量化单个细胞中的基因表达,并阐明复杂组织的分子和细胞构建模块。我们开发了一种新的贝叶斯层次模型称为细胞潜在狄利克雷分配(Celda)进行共聚类的基因转录模块和细胞亚群。Celda可以量化每个基因对每个模块的概率贡献,每个模块对每个细胞群的概率贡献以及每个细胞群对每个样品的概率贡献。在外周血单核细胞数据集中,Celda确定了增殖T细胞和浆细胞的亚群,这两个亚群被其他两个常见的单细胞工作流程遗漏。塞尔达还确定了转录模块,可用于表征跨细胞类型的独特和共享的生物程序。最后,Celda在模拟数据上将基因聚类成模块的性能优于其他方法。Celda提出了一种新的方法,用于表征scRNA-seq数据中的转录程序和细胞异质性。
Single-cell RNA-seq (scRNA-seq) has emerged as a powerful technique to quantify gene expression in individual cells and to elucidate the molecular and cellular building blocks of complex tissues. We developed a novel Bayesian hierarchical model called Cellular Latent Dirichlet Allocation (Celda) to perform co-clustering of genes into transcriptional modules and cells into subpopulations. Celda can quantify the probabilistic contribution of each gene to each module, each module to each cell population and each cell population to each sample. In a peripheral blood mononuclear cell dataset, Celda identified a subpopulation of proliferating T cells and a plasma cell which were missed by two other common single-cell workflows. Celda also identified transcriptional modules that could be used to characterize unique and shared biological programs across cell types. Finally, Celda outperformed other approaches for clustering genes into modules on simulated data. Celda presents a novel method for characterizing transcriptional programs and cellular heterogeneity in scRNA-seq data.