Model-based deep embedding for constrained clustering analysis of single cell RNA-seq data.

Model-based deep embedding for constrained clustering analysis of single cell RNA-seq data.
复制标题

DOI:
10.1038/s41467-021-22008-3
复制
发表时间:
2021-03-25
影响因子:
16.6
通讯作者:
Hakonarson H
Hakonarson H
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Tian T;Zhang J;Lin X;Wei Z;Hakonarson H

文献摘要

参考文献

被引文献

相似文献

聚类是基于单细胞的研究中的关键步骤。大多数现有的方法支持无监督聚类,没有任何领域知识的先验利用。当遇到scRNA-Seq数据的高维性和普遍的丢失事件时,纯粹无监督的聚类方法可能无法产生生物学上可解释的聚类,这使细胞类型分配复杂化。在这种情况下,唯一的办法是用户手动反复调整聚类参数,直到找到可接受的聚类。因此,获得有生物意义的聚类的路径可能是临时的和费力的。在这里,我们报告了一个原则性的聚类方法命名为scDCC,集成领域知识的聚类步骤。在数千到数万个细胞的各种scRNA-seq数据集上的实验表明,scDCC可以显著提高聚类性能,促进聚类和下游分析的可解释性,例如细胞类型分配。基于基因表达的相似性对细胞进行聚类是在scRNASeq数据中识别细胞类型的第一步。在这里,作者将生物学知识结合到聚类步骤中,以促进聚类的生物学可解释性,以及随后的细胞类型鉴定。
Clustering is a critical step in single cell-based studies. Most existing methods support unsupervised clustering without the a priori exploitation of any domain knowledge. When confronted by the high dimensionality and pervasive dropout events of scRNA-Seq data, purely unsupervised clustering methods may not produce biologically interpretable clusters, which complicates cell type assignment. In such cases, the only recourse is for the user to manually and repeatedly tweak clustering parameters until acceptable clusters are found. Consequently, the path to obtaining biologically meaningful clusters can be ad hoc and laborious. Here we report a principled clustering method named scDCC, that integrates domain knowledge into the clustering step. Experiments on various scRNA-seq datasets from thousands to tens of thousands of cells show that scDCC can significantly improve clustering performance, facilitating the interpretability of clusters and downstream analyses, such as cell type assignment. Clustering cells based on similarities in gene expression is the first step towards identifying cell types in scRNASeq data. Here the authors incorporate biological knowledge into the clustering step to facilitate the biological interpretability of clusters, and subsequent cell type identification.
DOI: 10.1016/j.cell.2015.05.047
发表时间: 2015-07-02
期刊: Cell
影响因子: 64.5
作者:
Levine JH;Simonds EF;Bendall SC;Davis KL;Amir el-AD;Tadmor MD;Litvin O;Fienberg HG;Jager A;Zunder ER;Finck R;Gedman AL;Radtke I;Downing JR;Pe'er D;Nolan GP
通讯作者: Nolan GP
DOI: 10.1038/s41467-018-06318-7
发表时间: 2018-10-22
影响因子: 16.6
作者:
MacParland SA;Liu JC;Ma XZ;Innes BT;Bartczak AM;Gage BK;Manuel J;Khuu N;Echeverri J;Linares I;Gupta R;Cheng ML;Liu LY;Camat D;Chung SW;Seliga RK;Shao Z;Lee E;Ogawa S;Ogawa M;Wilson MD;Fish JE;Selzner M;Ghanekar A;Grant D;Greig P;Sapisochin G;Selzner N;Winegarden N;Adeyi O;Keller G;Bader GD;McGilvray ID
通讯作者: McGilvray ID
DOI: 10.2307/2284239
发表时间: 1971-01-01
影响因子: 3.7
作者:
RAND, WM
通讯作者: RAND, WM
DOI: 10.1038/nmeth.4236
发表时间: 2017-05-01
期刊: NATURE METHODS
影响因子: 48
作者:
Kiselev, Vladimir Yu;Kirschner, Kristina;Hemberg, Martin
通讯作者: Hemberg, Martin
DOI: 10.1093/nar/gkw430
发表时间: 2016-07-27
影响因子: 14.9
作者:
Ji Z;Ji H
通讯作者: Ji H