Clustering methods for single-cell RNA-sequencing expression data: performance evaluation with varying sample sizes and cell compositions

Clustering methods for single-cell RNA-sequencing expression data: performance evaluation with varying sample sizes and cell compositions
复制标题

DOI:
10.1515/sagmb-2019-0004
复制
发表时间:
2019-10-01
影响因子:
0.9
通讯作者:
Suner, Asli
Suner, Asli
中科院分区:
数学4区
文献类型:
--
作者:
Suner, Asli

文献摘要

被引文献

相似文献

到目前为止,已经开发了许多专门的聚类方法来准确分析单细胞RNA测序(scRNA-seq)表达数据,并且已经发表了几个报告,记录了这些聚类方法在不同条件下的性能测量。然而,到目前为止,还没有关于在考虑到给定scRNA-seq数据集的样本大小和细胞组成的情况下对聚类方法的性能衡量进行系统评估的现有研究。在这里,使用具有已知样本大小和亚群数量以及不同水平的转录组复杂性的合成数据集,对11种选定的scRNA-seq聚类方法进行了全面的性能评估研究。结果表明,所研究的聚类方法的整体性能高度依赖于scRNA-seq数据集的样本大小和复杂性。在大多数情况下,随着给定表达数据集中的细胞数量的增加,获得了更好的聚类性能。这项研究的结果还强调了样本大小对于使用适当的聚类工具成功检测稀有细胞亚群的重要性。
A number of specialized clustering methods have been developed so far for the accurate analysis of single-cell RNA-sequencing (scRNA-seq) expression data, and several reports have been published documenting the performance measures of these clustering methods under different conditions. However, to date, there are no available studies regarding the systematic evaluation of the performance measures of the clustering methods taking into consideration the sample size and cell composition of a given scRNA-seq dataset. Herein, a comprehensive performance evaluation study of 11 selected scRNA-seq clustering methods was performed using synthetic datasets with known sample sizes and number of subpopulations, as well as varying levels of transcriptome complexity. The results indicate that the overall performance of the clustering methods under study are highly dependent on the sample size and complexity of the scRNA-seq dataset. In most of the cases, better clustering performances were obtained as the number of cells in a given expression dataset was increased. The findings of this study also highlight the importance of sample size for the successful detection of rare cell subpopulations with an appropriate clustering tool.