Dirichlet process mixture models for single-cell RNA-seq clustering.

Dirichlet process mixture models for single-cell RNA-seq clustering.
复制标题

DOI:
10.1242/bio.059001
复制
发表时间:
2022-04-15
期刊:
影响因子:
2.4
通讯作者:
--
中科院分区:
生物学4区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

基于基因表达的细胞聚类是单细胞 RNA 测序 (scRNA-seq) 数据分析的主要步骤之一。聚类分析的一个关键挑战是聚类数量未知,对于这个问题,仍然没有全面的解决方案。为了增强定义有意义的聚类分辨率的过程,我们在 scRNA-seq 数据聚类的背景下,将贝叶斯潜在狄利克雷分配 (LDA) 方法与其非参数对应的分层狄利克雷过程 (HDP) 进行比较。 HDP 的一个潜在主要优点是它不需要用户将簇的数量作为输入参数。虽然LDA已用于单细胞数据分析,但尚未与HDP进行详细比较。在这里,我们使用四个 scRNA-seq 数据集(免疫细胞、肾脏、胰腺和蜕膜/胎盘)比较 LDA 和 HDP 的细胞聚类性能,特别关注聚类数量。使用内在(DB-index)和外在(ARI)集群质量度量,我们表明LDA和HDP的性能依赖于数据集。我们描述了一种情况,与一系列具有不同数量簇的 LDA 聚类中的最佳性能相比,HDP 生成了更合适的聚类。然而,我们也观察到性能最佳的 LDA 簇数适当地捕获了主要生物学特征,而 HDP 往往会夸大簇数。总的来说,我们的研究强调了在分析 scRNA-seq 数据时仔细评估簇数量的重要性。摘要:狄利克雷混合模型(LDA 和 HDP)应用于 scRNA-Seq 数据中的细胞聚类。这里我们对scRNA-seq数据基于LDA和HDP模型的聚类进行了全面的比较。
Clustering of cells based on gene expression is one of the major steps in single-cell RNA-sequencing (scRNA-seq) data analysis. One key challenge in cluster analysis is the unknown number of clusters and, for this issue, there is still no comprehensive solution. To enhance the process of defining meaningful cluster resolution, we compare Bayesian latent Dirichlet allocation (LDA) method to its non-parametric counterpart, hierarchical Dirichlet process (HDP) in the context of clustering scRNA-seq data. A potential main advantage of HDP is that it does not require the number of clusters as an input parameter from the user. While LDA has been used in single-cell data analysis, it has not been compared in detail with HDP. Here, we compare the cell clustering performance of LDA and HDP using four scRNA-seq datasets (immune cells, kidney, pancreas and decidua/placenta), with a specific focus on cluster numbers. Using both intrinsic (DB-index) and extrinsic (ARI) cluster quality measures, we show that the performance of LDA and HDP is dataset dependent. We describe a case where HDP produced a more appropriate clustering compared to the best performer from a series of LDA clusterings with different numbers of clusters. However, we also observed cases where the best performing LDA cluster numbers appropriately capture the main biological features while HDP tended to inflate the number of clusters. Overall, our study highlights the importance of carefully assessing the number of clusters when analyzing scRNA-seq data. Summary: Dirichlet mixture models (LDA and HDP) are applied for clustering cells in scRNA-Seq data. Here we made a comprehensive comparison of LDA and HDP model-based clustering for scRNA-seq data.
DOI: 10.1038/s41586-019-0969-x
发表时间: 2019-02-28
期刊: NATURE
影响因子: 64.8
作者:
Cao, Junyue;Spielmann, Malte;Shendure, Jay
通讯作者: Shendure, Jay
DOI: 10.2307/2284239
发表时间: 1971-01-01
影响因子: 3.7
作者:
RAND, WM
通讯作者: RAND, WM
DOI: 10.1371/journal.pgen.1006599
发表时间: 2017-03-01
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Dey, Kushal K.;Hsiao, Chiaowen Joyce;Stephens, Matthew
通讯作者: Stephens, Matthew
DOI: 10.1504/ijcat.2013.052293
发表时间: 2013-01-01
影响因子: 1.1
作者:
Singh, Chandan Deep;Madan, Jatinder;Singh, Amrik
通讯作者: Singh, Amrik
DOI: 10.1162/jmlr.2003.3.4-5.993
发表时间: 2003-05-15
影响因子: 6
作者:
Blei, DM;Ng, AY;Jordan, MI
通讯作者: Jordan, MI