Demystifying "drop-outs" in single-cell UMI data
Demystifying "drop-outs" in single-cell UMI data
复制标题
DOI:
10.1186/s13059-020-02096-y
复制
发表时间:
2020-08-06
期刊:
影响因子:
12.3
通讯作者:
Chen, Mengjie
中科院分区:
文献类型:
--
作者:
Kim, Tae Hyun;Zhou, Xiang;Chen, Mengjie
Many existing pipelines for scRNA-seq data apply pre-processing steps such as normalization or imputation to account for excessive zeros or "drop-outs." Here, we extensively analyze diverse UMI data sets to show that clustering should be the foremost step of the workflow. We observe that most drop-outs disappear once cell-type heterogeneity is resolved, while imputing or normalizing heterogeneous data can introduce unwanted noise. We propose a novel framework HIPPO (Heterogeneity-Inspired Pre-Processing tOol) that leverages zero proportions to explain cellular heterogeneity and integrates feature selection with iterative clustering. HIPPO leads to downstream analysis with greater flexibility and interpretability compared to alternatives.