CyTOF workflow: differential discovery in high-throughput high-dimensional cytometry datasets.

CyTOF workflow: differential discovery in high-throughput high-dimensional cytometry datasets.
复制标题

DOI:
10.12688/f1000research.11622.1
复制
发表时间:
2017-01-01
期刊:
影响因子:
--
通讯作者:
Robinson, Mark D
Robinson, Mark D
中科院分区:
其他
文献类型:
--
作者:
Nowicka, Malgorzata;Krieg, Carsten;Robinson, Mark D

文献摘要

被引文献

相似文献

高维质量和流式细胞术(HDCyto)实验已成为高通量询问和表征细胞群体的首选方法。在这里,我们提出了一个基于R的HDCyto数据差异分析管道,主要基于Bioconductor软件包。我们计算定义细胞群体使用FlowSOM聚类,并促进一个可选的,但可重复的策略,手动合并算法生成的集群。我们的工作流程提供了不同的分析路径,包括细胞类型丰度与表型的关联或特定亚群内信号标记物的变化,或聚合信号的差异分析。重要的是,我们显示的差异分析是基于回归框架,其中HDCyto数据是响应;因此,我们能够模拟任意的实验设计,例如具有批效应的设计、配对设计等。特别地,我们将广义线性混合模型应用于细胞群体丰度的分析或信号标志物的细胞群体特异性分析,允许适当地对样本上的细胞计数或聚集信号的过度分散进行建模。为了支持正式的统计分析,我们鼓励在每一步进行探索性数据分析,包括质量控制(例如多维标度图)、聚类结果报告(降维、热图和树状图)和差异分析(例如聚合信号图)。
High dimensional mass and flow cytometry (HDCyto) experiments have become a method of choice for high throughput interrogation and characterization of cell populations.Here, we present an R-based pipeline for differential analyses of HDCyto data, largely based on Bioconductor packages. We computationally define cell populations using FlowSOM clustering, and facilitate an optional but reproducible strategy for manual merging of algorithm-generated clusters. Our workflow offers different analysis paths, including association of cell type abundance with a phenotype or changes in signaling markers within specific subpopulations, or differential analyses of aggregated signals. Importantly, the differential analyses we show are based on regression frameworks where the HDCyto data is the response; thus, we are able to model arbitrary experimental designs, such as those with batch effects, paired designs and so on. In particular, we apply generalized linear mixed models to analyses of cell population abundance or cell-population-specific analyses of signaling markers, allowing overdispersion in cell count or aggregated signals across samples to be appropriately modeled. To support the formal statistical analyses, we encourage exploratory data analysis at every step, including quality control (e.g. multi-dimensional scaling plots), reporting of clustering results (dimensionality reduction, heatmaps with dendrograms) and differential analyses (e.g. plots of aggregated signals).