DIMM-SC: a Dirichlet mixture model for clustering droplet-based single cell transcriptomic data

DIMM-SC: a Dirichlet mixture model for clustering droplet-based single cell transcriptomic data
复制标题

DIMM-SC:用于聚类基于液滴的单细胞转录组数据的狄利克雷混合模型

DOI:
10.1093/bioinformatics/btx490
复制
发表时间:
2018-01-01
期刊:
影响因子:
5.8
通讯作者:
Chen, Wei
Chen, Wei
中科院分区:
生物学3区
文献类型:
--
作者:
Sun, Zhe;Wang, Ting;Chen, Wei

文献摘要

被引文献

相似文献

动机 单细胞转录组测序(scRNA-Seq)已成为以单细胞分辨率研究细胞和分子过程的革命性工具。在现有技术中,最近开发的基于液滴的平台能够有效地并行处理数千个单细胞,并使用唯一分子标识符(UMI)直接计数转录本拷贝。尽管技术进步,统计方法和计算工具仍然缺乏分析基于液滴的scRNA-Seq数据。特别是,用于聚类大规模单细胞转录组学数据的基于模型的方法仍然未被探索。 结果 我们开发了DIMM-SC,一种用于聚类基于液滴的单细胞转录组数据的Dirichlet混合模型。这种方法明确地对来自scRNA-Seq实验的UMI计数数据进行建模,并通过Dirichlet混合先验来表征不同细胞簇之间的变化。我们进行了全面的模拟评估DIMM-SC,并将其与现有的聚类方法,如K-means,CellTree和Seurat进行比较。此外,我们分析了具有已知聚类标签的公共scRNA-Seq数据集和来自系统性硬化症研究的具有先前生物学知识的内部scRNA-Seq数据集,以基准测试和验证DIMM-SC。模拟研究和真实的数据应用都表明,与其他现有的聚类方法相比,DIMM-SC实现了显著提高的聚类准确性和更低的聚类变异性。更重要的是,作为一种基于模型的方法,DIMM-SC能够量化每个单细胞的聚类不确定性,促进严格的统计推断和生物学解释,这通常是现有聚类方法无法实现的。 可用性和实施 DIMM-SC已经在一个用户友好的R软件包中实现,详细的教程可以在www.pitt.edu/softwec47/erceell.html上找到。 接触 wei. chp.edu或hum@ccf. org。 补充资料 补充数据可在Bioinformatics在线获得。
Motivation Single cell transcriptome sequencing (scRNA-Seq) has become a revolutionary tool to study cellular and molecular processes at single cell resolution. Among existing technologies, the recently developed droplet-based platform enables efficient parallel processing of thousands of single cells with direct counting of transcript copies using Unique Molecular Identifier (UMI). Despite the technology advances, statistical methods and computational tools are still lacking for analyzing droplet-based scRNA-Seq data. Particularly, model-based approaches for clustering large-scale single cell transcriptomic data are still under-explored. Results We developed DIMM-SC, a Dirichlet Mixture Model for clustering droplet-based Single Cell transcriptomic data. This approach explicitly models UMI count data from scRNA-Seq experiments and characterizes variations across different cell clusters via a Dirichlet mixture prior. We performed comprehensive simulations to evaluate DIMM-SC and compared it with existing clustering methods such as K-means, CellTree and Seurat. In addition, we analyzed public scRNA-Seq datasets with known cluster labels and in-house scRNA-Seq datasets from a study of systemic sclerosis with prior biological knowledge to benchmark and validate DIMM-SC. Both simulation studies and real data applications demonstrated that overall, DIMM-SC achieves substantially improved clustering accuracy and much lower clustering variability compared to other existing clustering methods. More importantly, as a model-based approach, DIMM-SC is able to quantify the clustering uncertainty for each single cell, facilitating rigorous statistical inference and biological interpretations, which are typically unavailable from existing clustering methods. Availability and implementation DIMM-SC has been implemented in a user-friendly R package with a detailed tutorial available on www.pitt.edu/∼wec47/singlecell.html. Contact wei.chen@chp.edu or hum@ccf.org. Supplementary information Supplementary data are available at Bioinformatics online.