Hierarchical modeling for rare event detection and cell subset alignment across flow cytometry samples.

Hierarchical modeling for rare event detection and cell subset alignment across flow cytometry samples.
复制标题

DOI:
10.1371/journal.pcbi.1003130
复制
发表时间:
2013
影响因子:
4.3
通讯作者:
Chan C
Chan C
中科院分区:
生物学2区
文献类型:
--
作者:
Cron A;Gouttefangeas C;Frelinger J;Lin L;Singh SK;Britten CM;Welters MJ;van der Burg SH;West M;Chan C

文献摘要

参考文献

被引文献

相似文献

流式细胞术是多参数单细胞分析的典型方法,在疫苗和生物标志物研究中对于抗原特异性淋巴细胞的计数是必不可少的,这些淋巴细胞通常以极低的频率(0.1%或更少)被发现。流式细胞术数据的标准分析依赖于专家对细胞亚群的视觉识别,这是一个主观的过程,通常难以重现。另一种更客观的方法是使用统计模型以自动化的方式识别感兴趣的单元子集。自动化分析面临的两个具体挑战是检测极低频事件子集,而不会因预处理富集而使估计产生偏差,以及跨多个数据样本对齐细胞子集以进行比较分析的能力。在本文中,我们开发了对Dirichlet过程高斯混合模型(DPGMM)方法的分层建模扩展,我们之前描述了用于细胞子集识别的方法,并表明分层DPGMM (HDPGMM)自然地生成了一个对齐的数据模型,该模型捕获了多个样本的共性和差异。HDPGMM还通过同时分析多个样品共享信息来提高对极低频事件的灵敏度。我们在已知抗原特异性T细胞频率的临床相关参考外周血单个核细胞(PBMC)样本上验证了HDPGMM估计抗原特异性T细胞的准确性和可重复性。这些细胞样本利用逆转录病毒tcr转导的T细胞加入到自体PBMC样本中,通过hla -肽多聚体结合提供确定数量的抗原特异性T细胞。我们提供开源软件,可以利用多处理器和gpu加速来执行数值要求高的计算。我们表明,分层建模是一种有用的概率方法,可以提供细胞亚群的一致标记,并在定量抗原特异性免疫反应的背景下增加罕见事件检测的灵敏度。使用流式细胞术计数抗原特异性T细胞对于疫苗开发、免疫疗法监测和免疫生物标志物发现至关重要。对这些数据的分析具有挑战性,因为抗原特异性细胞的存在频率通常低于1000个外周血单核细胞(PBMC)中的1个。流式细胞术数据的标准分析依赖于专家对细胞亚群的视觉识别,这是一个主观的过程,通常难以重现。因此,人们对细胞子集识别的自动化方法有着浓厚的兴趣。这类自动化方法中最流行的一类是使用统计混合模型。我们提出了统计混合模型的层次扩展,它比标准混合模型有两个优点。首先,它提高了检测多个样本中存在的极其罕见事件集群的能力。其次,它可以通过在多个样本中以自然的方式排列分层公式产生的集群来直接比较细胞子集。我们在临床相关的参考PBMC样本上展示了该算法,这些样本具有已知频率的CD8 T细胞,用于表达癌睾丸抗原特异性T细胞受体(NY-ESO-1),并将其性能与其他流行的自动分析方法进行了比较。
Flow cytometry is the prototypical assay for multi-parameter single cell analysis, and is essential in vaccine and biomarker research for the enumeration of antigen-specific lymphocytes that are often found in extremely low frequencies (0.1% or less). Standard analysis of flow cytometry data relies on visual identification of cell subsets by experts, a process that is subjective and often difficult to reproduce. An alternative and more objective approach is the use of statistical models to identify cell subsets of interest in an automated fashion. Two specific challenges for automated analysis are to detect extremely low frequency event subsets without biasing the estimate by pre-processing enrichment, and the ability to align cell subsets across multiple data samples for comparative analysis. In this manuscript, we develop hierarchical modeling extensions to the Dirichlet Process Gaussian Mixture Model (DPGMM) approach we have previously described for cell subset identification, and show that the hierarchical DPGMM (HDPGMM) naturally generates an aligned data model that captures both commonalities and variations across multiple samples. HDPGMM also increases the sensitivity to extremely low frequency events by sharing information across multiple samples analyzed simultaneously. We validate the accuracy and reproducibility of HDPGMM estimates of antigen-specific T cells on clinically relevant reference peripheral blood mononuclear cell (PBMC) samples with known frequencies of antigen-specific T cells. These cell samples take advantage of retrovirally TCR-transduced T cells spiked into autologous PBMC samples to give a defined number of antigen-specific T cells detectable by HLA-peptide multimer binding. We provide open source software that can take advantage of both multiple processors and GPU-acceleration to perform the numerically-demanding computations. We show that hierarchical modeling is a useful probabilistic approach that can provide a consistent labeling of cell subsets and increase the sensitivity of rare event detection in the context of quantifying antigen-specific immune responses. The use of flow cytometry to count antigen-specific T cells is essential for vaccine development, monitoring of immune-based therapies and immune biomarker discovery. Analysis of such data is challenging because antigen-specific cells are often present in frequencies of less than 1 in 1,000 peripheral blood mononuclear cells (PBMC). Standard analysis of flow cytometry data relies on visual identification of cell subsets by experts, a process that is subjective and often difficult to reproduce. Consequently, there is intense interest in automated approaches for cell subset identification. One popular class of such automated approaches is the use of statistical mixture models. We propose a hierarchical extension of statistical mixture models that has two advantages over standard mixture models. First, it increases the ability to detect extremely rare event clusters that are present in multiple samples. Second, it enables direct comparison of cell subsets by aligning clusters across multiple samples in a natural way arising from the hierarchical formulation. We demonstrate the algorithm on clinically relevant reference PBMC samples with known frequencies of CD8 T cells engineered to express T cell receptors specific for the cancer-testis antigen (NY-ESO-1) and compare its performance with other popular automated analysis approaches.
DOI: 10.1155/2009/247646
发表时间: 2009
影响因子: --
作者:
Finak G;Bashashati A;Brinkman R;Gottardo R
通讯作者: Gottardo R
DOI: 10.1198/016214506000000302
发表时间: 2006-12-01
影响因子: 3.7
作者:
Teh, Yee Whye;Jordan, Michael I.;Blei, David M.
通讯作者: Blei, David M.
DOI: 10.1038/nmeth.2365
发表时间: 2013-03
期刊: Nature methods
影响因子: 48
作者:
Aghaeepour N;Finak G;FlowCAP Consortium;DREAM Consortium;Hoos H;Mosmann TR;Brinkman R;Gottardo R;Scheuermann RH
通讯作者: Scheuermann RH
DOI: 10.1198/016214501750332758
发表时间: 2001-03-01
影响因子: 3.7
作者:
Ishwaran, H;James, LF
通讯作者: James, LF
DOI: 10.1111/j.1467-9868.2004.05564.x
发表时间: 2004-01-01
影响因子: 5.8
作者:
Müller, P;Quintana, F;Rosner, G
通讯作者: Rosner, G