netNMF-sc: leveraging gene-gene interactions for imputation and dimensionality reduction in single-cell expression analysis

netNMF-sc: leveraging gene-gene interactions for imputation and dimensionality reduction in single-cell expression analysis
复制标题

DOI:
10.1101/gr.251603.119
复制
发表时间:
2020-02-01
期刊:
影响因子:
7
通讯作者:
Raphael, Benjamin J.
Raphael, Benjamin J.
中科院分区:
生物学1区
文献类型:
--
作者:
Elyanow, Rebecca;Dumitrascu, Bianca;Raphael, Benjamin J.

文献摘要

被引文献

相似文献

单细胞RNA测序(scRNA-seq)使高通量测量单细胞中的RNA表达成为可能。然而,由于技术限制,scRNA-seq数据通常包含单个细胞中许多转录本的零计数。这些零计数或辍学事件使使用为批量RNA-SEQ数据开发的标准方法分析scRNA-SEQ数据变得复杂。当前的scRNA-seq分析方法通常通过在较低维空间中组合跨细胞的信息来克服丢失,利用细胞通常占据少量RNA表达状态的观察。我们介绍了netNMF-sc,这是一种利用细胞和基因之间的信息进行scRNA-seq分析的算法。NetNMF-sc使用网络正则化的非负矩阵分解来学习scRNA-seq转录计数的低维表示。网络正则化利用了基因-基因相互作用的先验知识,鼓励具有已知相互作用的基因对在低维表示中彼此接近。由此产生的矩阵因式分解计算了零计数和非零计数的基因丰度,并可用于将细胞聚集成有意义的亚群。我们表明,netNMF-sc在使用模拟和真实scRNA-seq数据对细胞进行聚类和估计基因-基因协方差方面优于现有方法,并且在较高的辍学率(例如,60%)下具有越来越大的优势。我们还表明,netNMF-sc的结果对输入网络的变化是健壮的,更具代表性的网络导致更大的性能收益。
Single-cell RNA-sequencing (scRNA-seq) enables high-throughput measurement of RNA expression in single cells. However, because of technical limitations, scRNA-seq data often contain zero counts for many transcripts in individual cells. These zero counts, or dropout events, complicate the analysis of scRNA-seq data using standard methods developed for bulk RNA-seq data. Current scRNA-seq analysis methods typically overcome dropout by combining information across cells in a lower-dimensional space, leveraging the observation that cells generally occupy a small number of RNA expression states. We introduce netNMF-sc, an algorithm for scRNA-seq analysis that leverages information across both cells and genes. netNMF-sc learns a low-dimensional representation of scRNA-seq transcript counts using network-regularized non-negative matrix factorization. The network regularization takes advantage of prior knowledge of gene-gene interactions, encouraging pairs of genes with known interactions to be nearby each other in the low-dimensional representation. The resulting matrix factorization imputes gene abundance for both zero and nonzero counts and can be used to cluster cells into meaningful subpopulations. We show that netNMF-sc outperforms existing methods at clustering cells and estimating gene-gene covariance using both simulated and real scRNA-seq data, with increasing advantages at higher dropout rates (e.g., >60%). We also show that the results from netNMF-sc are robust to variation in the input network, with more representative networks leading to greater performance gains.