Predicting Cell Populations in Single Cell Mass Cytometry Data

Predicting Cell Populations in Single Cell Mass Cytometry Data
复制标题

DOI:
10.1002/cyto.a.23738
复制
发表时间:
2019-07-01
期刊:
影响因子:
3.7
通讯作者:
Mahfouz, Ahmed
Mahfouz, Ahmed
中科院分区:
生物学4区
文献类型:
--
作者:
Abdelaal, Tamim;van Unen, Vincent;Mahfouz, Ahmed

文献摘要

被引文献

相似文献

飞行时间质谱仪(CyTOF)是一种在单细胞水平进行高维分析的有价值的技术。不同细胞群的鉴定是数据分析过程中的一项重要任务。许多聚类工具可以执行此任务,这对于在探索性实验中识别“新”细胞群至关重要。然而,依赖于聚类是费力的,因为它通常涉及手动注释,这显著限制了在不同样品中鉴定细胞群体的再现性。后者在比较不同条件的研究中特别重要,例如在队列研究中。从一组注释的细胞中学习细胞群解决了这些问题。然而,目前可用的用于自动细胞群体识别的方法要么是复杂的,依赖于在学习过程中关于群体的先前生物学知识,要么只能识别典型的细胞群体。我们建议使用线性判别分析(LDA)分类器自动识别细胞飞行时间数据中的细胞群。LDA在四个基准数据集上的性能优于两种最先进的算法。与更复杂的分类器相比,LDA在可解释的性能,可重复性和可扩展性方面具有很大的优势,可以扩展到具有更深注释的更大数据集。我们将LDA应用于类似于350万个细胞的数据集,这些细胞代表人类粘液免疫系统中的57个细胞群体。LDA对丰富的细胞群体以及大多数稀有细胞群体具有高性能,并提供细胞群体频率的准确估计。基于估计的后验概率,进一步结合拒绝选项,允许LDA识别在训练期间未遇到的先前未知的(新的)细胞群。总而言之,使用LDA对细胞群体组成的可重复预测开辟了基于CyTOF数据分析大型队列研究的可能性。(C)2019年,任作家。Cytometry Part A由Wiley Periodicals,Inc.出版。
Mass cytometry by time-of-flight (CyTOF) is a valuable technology for high-dimensional analysis at the single cell level. Identification of different cell populations is an important task during the data analysis. Many clustering tools can perform this task, which is essential to identify "new" cell populations in explorative experiments. However, relying on clustering is laborious since it often involves manual annotation, which significantly limits the reproducibility of identifying cell-populations across different samples. The latter is particularly important in studies comparing different conditions, for example in cohort studies. Learning cell populations from an annotated set of cells solves these problems. However, currently available methods for automatic cell population identification are either complex, dependent on prior biological knowledge about the populations during the learning process, or can only identify canonical cell populations. We propose to use a linear discriminant analysis (LDA) classifier to automatically identify cell populations in CyTOF data. LDA outperforms two state-of-the-art algorithms on four benchmark datasets. Compared to more complex classifiers, LDA has substantial advantages with respect to the interpretable performance, reproducibility, and scalability to larger datasets with deeper annotations. We apply LDA to a dataset of similar to 3.5 million cells representing 57 cell populations in the Human Mucosal Immune System. LDA has high performance on abundant cell populations as well as the majority of rare cell populations, and provides accurate estimates of cell population frequencies. Further incorporating a rejection option, based on the estimated posterior probabilities, allows LDA to identify previously unknown (new) cell populations that were not encountered during training. Altogether, reproducible prediction of cell population compositions using LDA opens up possibilities to analyze large cohort studies based on CyTOF data. (C) 2019 The Authors. Cytometry Part A published by Wiley Periodicals, Inc.