A comparison of automatic cell identification methods for single-cell RNA sequencing data

A comparison of automatic cell identification methods for single-cell RNA sequencing data
复制标题

DOI:
10.1186/s13059-019-1795-z
复制
发表时间:
2019-09-09
期刊:
影响因子:
12.3
通讯作者:
Mahfouz, Ahmed
Mahfouz, Ahmed
中科院分区:
生物学1区
文献类型:
--
作者:
Abdelaal, Tamim;Michielsen, Lieke;Mahfouz, Ahmed

文献摘要

被引文献

相似文献

背景单细胞转录组学正在迅速推进我们对复杂组织和生物体的细胞组成的理解。大多数分析管道的一个主要限制是依赖手动注释来确定细胞身份,这是耗时且不可重现的。细胞和样本数量的指数增长促使了用于细胞自动识别的监督分类方法的适应和发展。结果在这里,我们对22种分类方法进行了基准测试,这些方法自动分配细胞身份,包括单细胞特定分类器和通用分类器。使用27个公开可用的单细胞RNA测序数据集评估了这些方法的性能,这些数据集具有不同的大小、技术、物种和复杂程度。我们使用两个实验设置来评估每种方法在基于准确度、未分类单元百分比和计算时间的数据集内预测(数据集内)和跨数据集预测(数据集间)的性能。我们进一步评估了这些方法对输入特征、每个种群的单元数的敏感性,以及它们在不同标注级别和数据集上的性能。我们发现,大多数分类器在各种数据集上都表现得很好,但对于具有重叠类或深层标注的复杂数据集,其准确率会降低。在不同的实验中,通用支持向量机分类器的性能总体上是最好的。结论我们对单细胞RNA测序数据的自动细胞识别方法进行了全面的评价。用于评估的所有代码都可以在GitHub()上找到。此外,我们提供了Snakemake工作流来促进基准测试,并支持新方法和新数据集的扩展。
Background Single-cell transcriptomics is rapidly advancing our understanding of the cellular composition of complex tissues and organisms. A major limitation in most analysis pipelines is the reliance on manual annotations to determine cell identities, which are time-consuming and irreproducible. The exponential growth in the number of cells and samples has prompted the adaptation and development of supervised classification methods for automatic cell identification. Results Here, we benchmarked 22 classification methods that automatically assign cell identities including single-cell-specific and general-purpose classifiers. The performance of the methods is evaluated using 27 publicly available single-cell RNA sequencing datasets of different sizes, technologies, species, and levels of complexity. We use 2 experimental setups to evaluate the performance of each method for within dataset predictions (intra-dataset) and across datasets (inter-dataset) based on accuracy, percentage of unclassified cells, and computation time. We further evaluate the methods' sensitivity to the input features, number of cells per population, and their performance across different annotation levels and datasets. We find that most classifiers perform well on a variety of datasets with decreased accuracy for complex datasets with overlapping classes or deep annotations. The general-purpose support vector machine classifier has overall the best performance across the different experiments. Conclusions We present a comprehensive evaluation of automatic cell identification methods for single-cell RNA sequencing data. All the code used for the evaluation is available on GitHub (). Additionally, we provide a Snakemake workflow to facilitate the benchmarking and to support the extension of new methods and new datasets.