Pooling across cells to normalize single-cell RNA sequencing data with many zero counts.

Pooling across cells to normalize single-cell RNA sequencing data with many zero counts.
复制标题

DOI:
10.1186/s13059-016-0947-7
复制
发表时间:
2016-04-27
期刊:
影响因子:
12.3
通讯作者:
Marioni JC
Marioni JC
中科院分区:
生物学1区
文献类型:
--
作者:
Lun AT;Bach K;Marioni JC

文献摘要

被引文献

相似文献

单细胞RNA测序数据的标准化对于在下游分析之前消除特定细胞的偏向是必要的。然而,这对于噪声较大的单元格数据来说并不简单,其中许多计数都为零。我们提出了一种新的方法,其中跨单元格池对表达式值进行求和,并将求和值用于归一化。然后对基于池的大小因子进行去卷积,以产生基于单元格的因子。我们的去卷积方法优于现有的在模拟数据中精确归一化特定细胞偏见的方法。在真实数据中也观察到了类似的行为,其中去卷积提高了下游分析结果的相关性。本文的在线版本(doi:10.1186/s13059-0160947-7)包含补充材料,授权用户可以使用。
Normalization of single-cell RNA sequencing data is necessary to eliminate cell-specific biases prior to downstream analyses. However, this is not straightforward for noisy single-cell data where many counts are zero. We present a novel approach where expression values are summed across pools of cells, and the summed values are used for normalization. Pool-based size factors are then deconvolved to yield cell-based factors. Our deconvolution approach outperforms existing methods for accurate normalization of cell-specific biases in simulated data. Similar behavior is observed in real data, where deconvolution improves the relevance of results of downstream analyses. The online version of this article (doi:10.1186/s13059-016-0947-7) contains supplementary material, which is available to authorized users.