CIPR: a web-based R/shiny app and R package to annotate cell clusters in single cell RNA sequencing experiments

CIPR: a web-based R/shiny app and R package to annotate cell clusters in single cell RNA sequencing experiments
复制标题

DOI:
10.1186/s12859-020-3538-2
复制
发表时间:
2020-05-15
期刊:
影响因子:
3
通讯作者:
O'Connell, Ryan M.
O'Connell, Ryan M.
中科院分区:
生物学4区
文献类型:
--
作者:
Ekiz, H. Atakan;Conley, Christopher J.;O'Connell, Ryan M.

文献摘要

被引文献

相似文献

背景 单细胞RNA测序(scRNAseq)为健康和疾病状态下的细胞异质性和功能状态提供了极有价值的见解。在scRNAseq数据的分析过程中,对细胞簇的生物学特性进行注释是下游分析之前的重要步骤,并且在技术上仍然具有挑战性。目前用于注释单细胞簇的解决方案通常缺乏图形用户界面,可能计算量较大,或者适用范围有限。另一方面,通过检查标记基因的表达来手动注释单细胞簇可能具有主观性且劳动强度大。为了提高scRNAseq数据中细胞簇注释的质量和效率,我们推出了一个基于网络的R/Shiny应用程序和R包,即簇特性预测器(CIPR),它提供了一个图形用户界面,可以针对小鼠或人类参考数据,或者用户提供的自定义数据集,快速对未知细胞簇的基因表达谱进行评分。CIPR可以很容易地整合到当前的流程中,以促进scRNAseq数据分析。 结果 CIPR采用多种方法在簇水平计算特性分数,并且可以接受流行的scRNAseq分析软件生成的输入。CIPR提供2个小鼠和5个人类参考数据集,其流程允许进行物种间比较,并能够上传自定义参考数据集用于专门研究。从分析中过滤掉低变异性基因以及排除不相关的参考细胞子集的选项可以提高CIPR的区分能力,这表明它可以针对不同的实验环境进行调整。将CIPR与现有的功能相似的软件进行基准测试表明,我们的算法对计算资源的需求较低,运行速度明显更快,并且在一项涉及肿瘤浸润免疫细胞的scRNAseq实验中,对多个细胞簇提供了准确的预测。 结论 CIPR通过以客观且高效的方式注释未知细胞簇,促进了scRNAseq数据分析。由于Shiny框架具有平台独立性,并且对编程经验要求极低,使得该软件能够被不同背景的研究人员使用。CIPR可以准确预测多种细胞簇的特性,并且可以在广泛的研究领域的各种实验环境中使用。
Background Single cell RNA sequencing (scRNAseq) has provided invaluable insights into cellular heterogeneity and functional states in health and disease. During the analysis of scRNAseq data, annotating the biological identity of cell clusters is an important step before downstream analyses and it remains technically challenging. The current solutions for annotating single cell clusters generally lack a graphical user interface, can be computationally intensive or have a limited scope. On the other hand, manually annotating single cell clusters by examining the expression of marker genes can be subjective and labor-intensive. To improve the quality and efficiency of annotating cell clusters in scRNAseq data, we present a web-based R/Shiny app and R package, Cluster Identity PRedictor (CIPR), which provides a graphical user interface to quickly score gene expression profiles of unknown cell clusters against mouse or human references, or a custom dataset provided by the user. CIPR can be easily integrated into the current pipelines to facilitate scRNAseq data analysis. Results CIPR employs multiple approaches for calculating the identity score at the cluster level and can accept inputs generated by popular scRNAseq analysis software. CIPR provides 2 mouse and 5 human reference datasets, and its pipeline allows inter-species comparisons and the ability to upload a custom reference dataset for specialized studies. The option to filter out lowly variable genes and to exclude irrelevant reference cell subsets from the analysis can improve the discriminatory power of CIPR suggesting that it can be tailored to different experimental contexts. Benchmarking CIPR against existing functionally similar software revealed that our algorithm is less computationally demanding, it performs significantly faster and provides accurate predictions for multiple cell clusters in a scRNAseq experiment involving tumor-infiltrating immune cells. Conclusions CIPR facilitates scRNAseq data analysis by annotating unknown cell clusters in an objective and efficient manner. Platform independence owing to Shiny framework and the requirement for a minimal programming experience allows this software to be used by researchers from different backgrounds. CIPR can accurately predict the identity of a variety of cell clusters and can be used in various experimental contexts across a broad spectrum of research areas.