Enhancing droplet-based single-nucleus RNA-seq resolution using the semi-supervised machine learning classifier DIEM

Enhancing droplet-based single-nucleus RNA-seq resolution using the semi-supervised machine learning classifier DIEM
复制标题

DOI:
10.1038/s41598-020-67513-5
复制
发表时间:
2020-07-03
期刊:
影响因子:
4.6
通讯作者:
Pajukanta, Paivi
Pajukanta, Paivi
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Alvarez, Marcus;Rahmani, Elior;Pajukanta, Paivi

文献摘要

被引文献

相似文献

单核RNA测序(snRNA-seq)测量单个细胞核而不是细胞中的基因表达,从而可以在实体组织中进行无偏倚的细胞类型表征。我们观察到snRNA-seq通常会受到大量环境RNA的污染,这可能导致下游分析出现偏差,例如如果被忽略,则会识别出虚假的细胞类型。我们提出了一种新的方法来量化snRNA-seq实验中的污染和过滤液滴,称为使用期望最大化(DIEM)的碎片识别。我们的基于似然的方法模型的碎片和细胞类型,这是使用EM估计的基因表达分布。我们使用三个snRNA-seq数据集评估了DIEM:(1)体外人分化前脂肪细胞,(2)新鲜小鼠脑组织和(3)来自六个个体的人冷冻脂肪组织(AT)。所有三个数据集都显示了细胞核RNA污染的证据,我们观察到现有的方法无法解释污染的液滴,并导致虚假的细胞类型。与使用这些最先进方法进行过滤相比,DIEM可以更好地去除含有高水平核外RNA的液滴,并产生更高质量的簇。尽管DIEM是为snRNA-seq设计的,但我们的聚类策略也成功地过滤了单细胞RNA-seq数据。总之,我们的新方法DIEM快速有效地从基于单细胞的数据中去除碎片污染的液滴,从而实现更清洁的下游分析。我们的代码可在https://github.com/marcalva/diem上免费获得。
Single-nucleus RNA sequencing (snRNA-seq) measures gene expression in individual nuclei instead of cells, allowing for unbiased cell type characterization in solid tissues. We observe that snRNA-seq is commonly subject to contamination by high amounts of ambient RNA, which can lead to biased downstream analyses, such as identification of spurious cell types if overlooked. We present a novel approach to quantify contamination and filter droplets in snRNA-seq experiments, called Debris Identification using Expectation Maximization (DIEM). Our likelihood-based approach models the gene expression distribution of debris and cell types, which are estimated using EM. We evaluated DIEM using three snRNA-seq data sets: (1) human differentiating preadipocytes in vitro, (2) fresh mouse brain tissue, and (3) human frozen adipose tissue (AT) from six individuals. All three data sets showed evidence of extranuclear RNA contamination, and we observed that existing methods fail to account for contaminated droplets and led to spurious cell types. When compared to filtering using these state of the art methods, DIEM better removed droplets containing high levels of extranuclear RNA and led to higher quality clusters. Although DIEM was designed for snRNA-seq, our clustering strategy also successfully filtered single-cell RNA-seq data. To conclude, our novel method DIEM removes debris-contaminated droplets from single-cell-based data fast and effectively, leading to cleaner downstream analysis. Our code is freely available for use at https://github.com/marcalva/diem.