A scalable SCENIC workflow for single-cell gene regulatory network analysis

A scalable SCENIC workflow for single-cell gene regulatory network analysis
复制标题

DOI:
10.1038/s41596-020-0336-2
复制
发表时间:
2020-06-19
期刊:
影响因子:
14.8
通讯作者:
Aerts, Stein
Aerts, Stein
中科院分区:
生物学1区
文献类型:
--
作者:
van de Sande, Bram;Flerin, Christopher;Aerts, Stein

文献摘要

被引文献

相似文献

SCENIC 是一个通过网络推理和基序富集来预测细胞类型特异性转录因子的计算管道。在这里,作者描述了 pySCENIC 的详细协议:Python 中更快、基于容器的实现。该协议解释了如何使用软件容器和 Nextflow 管道对单细胞 RNA 测序数据执行快速 SCENIC 分析以及标准最佳实践步骤。 SCENIC 重建调节子(即转录因子及其目标基因),评估单个细胞中这些发现的调节子的活性,并使用这些细胞活动模式来寻找有意义的细胞簇。在这里,我们展示了 SCENIC 的改进版本,具有多项进步。 SCENIC 已经用 Python (pySCENIC) 进行了重构和重新实现,速度提高了十倍,并且已打包到容器中以方便使用。现在还可以使用表观基因组轨迹数据库以及基序来完善调节子。在此协议中,我们解释了 SCENIC 的不同步骤:工作流程从描述所有细胞基因丰度的计数矩阵开始,由三个阶段组成。首先,使用每目标回归方法 (GRNBoost2) 推断共表达模块。接下来,使用顺式调控基序发现 (cisTarget) 从这些模块中修剪间接目标。最后,这些调节子的活性通过调节子目标基因 (AUCell) 的富集分数进行量化。非线性投影方法可用于根据这些调节子的细胞活动模式显示细胞的视觉分组。结果可以导出为 loom 文件并在 SCope Web 应用程序中可视化。该协议通过两个用例进行说明:外周血单核细胞数据集和一组单细胞 RNA 测序癌症实验。对于包含 10,000 个基因和 50,000 个细胞的数据集,管道运行时间为
SCENIC is a computational pipeline to predict cell-type-specific transcription factors through network inference and motif enrichment. Here the authors describe a detailed protocol for pySCENIC: a faster, container-based implementation in Python.This protocol explains how to perform a fast SCENIC analysis alongside standard best practices steps on single-cell RNA-sequencing data using software containers and Nextflow pipelines. SCENIC reconstructs regulons (i.e., transcription factors and their target genes) assesses the activity of these discovered regulons in individual cells and uses these cellular activity patterns to find meaningful clusters of cells. Here we present an improved version of SCENIC with several advances. SCENIC has been refactored and reimplemented in Python (pySCENIC), resulting in a tenfold increase in speed, and has been packaged into containers for ease of use. It is now also possible to use epigenomic track databases, as well as motifs, to refine regulons. In this protocol, we explain the different steps of SCENIC: the workflow starts from the count matrix depicting the gene abundances for all cells and consists of three stages. First, coexpression modules are inferred using a regression per-target approach (GRNBoost2). Next, the indirect targets are pruned from these modules using cis-regulatory motif discovery (cisTarget). Lastly, the activity of these regulons is quantified via an enrichment score for the regulon's target genes (AUCell). Nonlinear projection methods can be used to display visual groupings of cells based on the cellular activity patterns of these regulons. The results can be exported as a loom file and visualized in the SCope web application. This protocol is illustrated on two use cases: a peripheral blood mononuclear cell data set and a panel of single-cell RNA-sequencing cancer experiments. For a data set of 10,000 genes and 50,000 cells, the pipeline runs in