Demystifying "drop-outs" in single-cell UMI data

Demystifying "drop-outs" in single-cell UMI data
复制标题

DOI:
10.1186/s13059-020-02096-y
复制
发表时间:
2020-08-06
期刊:
影响因子:
12.3
通讯作者:
Chen, Mengjie
Chen, Mengjie
中科院分区:
生物学1区
文献类型:
--
作者:
Kim, Tae Hyun;Zhou, Xiang;Chen, Mengjie

文献摘要

被引文献

相似文献

scRNA-seq数据的许多现有管道应用预处理步骤,例如归一化或插补,以考虑过多的零或“脱落”。“在这里,我们广泛分析了不同的UMI数据集,以表明聚类应该是工作流程的最重要步骤。我们观察到,一旦细胞类型的异质性得到解决,大多数脱落消失,而输入或归一化异质性数据可能会引入不必要的噪音。我们提出了一个新的框架HIPPO(异质性启发的预处理工具),利用零比例来解释细胞异质性,并将特征选择与迭代聚类相结合。HIPPO导致下游分析具有更大的灵活性和可解释性相比,替代品。
Many existing pipelines for scRNA-seq data apply pre-processing steps such as normalization or imputation to account for excessive zeros or "drop-outs." Here, we extensively analyze diverse UMI data sets to show that clustering should be the foremost step of the workflow. We observe that most drop-outs disappear once cell-type heterogeneity is resolved, while imputing or normalizing heterogeneous data can introduce unwanted noise. We propose a novel framework HIPPO (Heterogeneity-Inspired Pre-Processing tOol) that leverages zero proportions to explain cellular heterogeneity and integrates feature selection with iterative clustering. HIPPO leads to downstream analysis with greater flexibility and interpretability compared to alternatives.