SECANT: a biology-guided semi-supervised method for clustering, classification, and annotation of single-cell multi-omics

SECANT: a biology-guided semi-supervised method for clustering, classification, and annotation of single-cell multi-omics
复制标题

DOI:
10.1101/2020.11.06.371849
复制
发表时间:
2020-11
期刊:
PNAS Nexus
影响因子:
--
通讯作者:
Xinjun Wang;Zhongli Xu;Xu Zhou;Yanfu Zhang;Heng Huang;Ying Ding;R. Duerr;Wei Chen
Xinjun Wang;Zhongli Xu;Xu Zhou;Yanfu Zhang;Heng Huang;Ying Ding;R. Duerr;Wei Chen
中科院分区:
其他
文献类型:
--
作者:
Xinjun Wang;Zhongli Xu;Xu Zhou;Yanfu Zhang;Heng Huang;Ying Ding;R. Duerr;Wei Chen

文献摘要

相似文献

单细胞测序(SCRNA-SEQ)技术的最新进展,例如通过测序(CITE-SEQ)对转录组和表位的细胞索引(CITE-SEQ)允许研究人员在单细胞分辨率下同时量化细胞表面蛋白丰度和RNA表达。尽管Cite-seq和其他类似技术迅速获得了巨大的流行,但分析这种新型单细胞多词数据数据的新方法仍在迫切需要。有限的可用工具利用数据驱动的方法,这可能会破坏表面蛋白数据的生物学重要性。在这项研究中,我们开发了SECANT,这是一种生物学指导的半监督方法,用于单细胞多矩的聚类,分类和注释。 SECANT可用于分析Cite-Seq数据,或者共同分析Cite-Seq和Scrna-Seq数据。 SECANT的新颖性包括1)使用从表面蛋白数据确定为细胞聚类指南的自信细胞类型标签,2)为每个细胞簇提供自信的细胞类型的一般注释,3)完全利用具有不确定或缺失的单元格标签的单元格到提高性能,4)准确预测从表面蛋白数据中确定的SCRNA-SEQ数据的自信细胞类型。此外,作为一种基于模型的方法,SECANT可以量化结果的不确定性,并且我们的框架可以轻松扩展以处理其他类型的多媒体数据。我们通过模拟研究以及对公共和内部实际数据集的分析成功证明了Secant的有效性和优势。我们认为,这种新方法将极大地帮助研究人员表征新颖的细胞类型,并使用单细胞多词数据进行新的生物学发现。
The recent advance of single cell sequencing (scRNA-seq) technology such as Cellular Indexing of Transcriptomes and Epitopes by Sequencing (CITE-seq) allows researchers to quantify cell surface protein abundance and RNA expression simultaneously at single cell resolution. Although CITE-seq and other similar technologies have quickly gained enormous popularity, novel methods for analyzing this new type of single cell multi-omics data are still in urgent need. A limited number of available tools utilize data-driven approach, which may undermine the biological importance of surface protein data. In this study, we developed SECANT, a biology-guided SEmi-supervised method for Clustering, classification, and ANnoTation of single-cell multi-omics. SECANT can be used to analyze CITE-seq data, or jointly analyze CITE-seq and scRNA-seq data. The novelties of SECANT include 1) using confident cell type labels identified from surface protein data as guidance for cell clustering, 2) providing general annotation of confident cell types for each cell cluster, 3) fully utilizing cells with uncertain or missing cell type labels to increase performance, and 4) accurate prediction of confident cell types identified from surface protein data for scRNA-seq data. Besides, as a model-based approach, SECANT can quantify the uncertainty of the results, and our framework can be easily extended to handle other types of multi-omics data. We successfully demonstrated the validity and advantages of SECANT via simulation studies and analysis of public and in-house real datasets. We believe this new method will greatly help researchers characterize novel cell types and make new biological discoveries using single cell multi-omics data.