Jointly defining cell types from multiple single-cell datasets using LIGER.

Jointly defining cell types from multiple single-cell datasets using LIGER.
复制标题

DOI:
10.1038/s41596-020-0391-8
复制
发表时间:
2020-11
期刊:
影响因子:
14.8
通讯作者:
Welch JD
Welch JD
中科院分区:
生物学1区
文献类型:
--
作者:
Liu J;Gao C;Sodicoff J;Kozareva V;Macosko EZ;Welch JD

文献摘要

参考文献

被引文献

相似文献

高通量单细胞测序技术在使用基因表达和表观基因组状态以无偏方式定义细胞类型方面具有巨大潜力。实现这一潜力的一个关键挑战是将来自多个协议、生物背景和数据模式的单细胞数据集整合到细胞身份的联合定义中。我们之前开发了一种名为基因组实验关系关联推理(LIGER,https://github.com/MacoskoLab/liger)的方法,该方法使用整合非负矩阵因子分解来解决这一挑战。在这里,我们提供了一个使用LIGER从多个单细胞数据集中联合定义细胞类型的分步协议。该协议的主要步骤包括数据预处理和归一化、联合因子分解、分位数归一化和联合聚类以及可视化。我们描述了如何从单细胞RNA-seq(scRNA-seq)和单核ATAC-seq(snATAC-seq)数据中联合定义细胞类型,但类似的步骤适用于广泛的其他设置和数据类型,包括跨物种分析,单核DNA甲基化和空间转录组学。我们的协议包含预期结果的示例,描述了常见的陷阱,并且仅依赖于我们免费提供的LIGER开源R实现。我们还提供Rmarkdown教程,展示每个代码段的输出。根据数据集的大小,分析过程可以在1-4小时内完成,并且假设没有专门的生物信息学培训。在这里,作者描述了使用基于R的软件工具LIGER整合来自不同实验或模式的单细胞测序数据集以识别常见和不同细胞类型的分步程序。
High-throughput single-cell sequencing technologies hold tremendous potential for defining cell types in an unbiased fashion using gene expression and epigenomic state. A key challenge in realizing this potential is integrating single-cell datasets from multiple protocols, biological contexts, and data modalities into a joint definition of cellular identity. We previously developed an approach called Linked Inference of Genomic Experimental Relationships (LIGER, https://github.com/MacoskoLab/liger) that uses integrative nonnegative matrix factorization to address this challenge. Here, we provide a step-by-step protocol for using LIGER to jointly define cell types from multiple single-cell datasets. The main steps of the protocol include data preprocessing and normalization, joint factorization, quantile normalization and joint clustering, and visualization. We describe how to jointly define cell types from single-cell RNA-seq (scRNA-seq) and single-nucleus ATAC-seq (snATAC-seq) data, but similar steps apply across a wide range of other settings and data types, including cross-species analysis, single-nucleus DNA methylation, and spatial transcriptomics. Our protocol contains examples of expected results, describes common pitfalls, and relies only on our freely available, open-source R implementation of LIGER. We also provide Rmarkdown tutorials showing the outputs from each individual code segment. The analysis process can be performed in 1–4 hours depending on dataset size and assumes no specialized bioinformatics training. Here, the authors describe step-by-step procedures for integrating single cell sequencing datasets from different experiments or modalities to identify common and distinct cell types using the R-based software tool LIGER.
DOI: 10.1016/j.cell.2018.07.028
发表时间: 2018-08-09
期刊: Cell
影响因子: 64.5
作者:
Saunders A;Macosko EZ;Wysoker A;Goldman M;Krienen FM;de Rivera H;Bien E;Baum M;Bortolin L;Wang S;Goeva A;Nemesh J;Kamitaki N;Brumbaugh S;Kulp D;McCarroll SA
通讯作者: McCarroll SA
DOI: 10.1038/s41592-019-0466-z
发表时间: 2019-08-01
期刊: NATURE METHODS
影响因子: 48
作者:
Barkas, Nikolas;Petukhov, Viktor;Kharchenko, Peter V.
通讯作者: Kharchenko, Peter V.
DOI: 10.1093/bioinformatics/btv544
发表时间: 2016-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Yang, Zi;Michailidis, George
通讯作者: Michailidis, George
DOI: 10.1038/s41592-019-0529-1
发表时间: 2019-10-01
期刊: NATURE METHODS
影响因子: 48
作者:
Zhang, Allen W.;O'Flanagan, Ciara;Shah, Sohrab P.
通讯作者: Shah, Sohrab P.
DOI: 10.1038/s41586-018-0393-7
发表时间: 2018-08
期刊: Nature
影响因子: 64.8
作者:
Montoro DT;Haber AL;Biton M;Vinarsky V;Lin B;Birket SE;Yuan F;Chen S;Leung HM;Villoria J;Rogel N;Burgin G;Tsankov AM;Waghray A;Slyper M;Waldman J;Nguyen L;Dionne D;Rozenblatt-Rosen O;Tata PR;Mou H;Shivaraju M;Bihler H;Mense M;Tearney GJ;Rowe SM;Engelhardt JF;Regev A;Rajagopal J
通讯作者: Rajagopal J