False signals induced by single-cell imputation.

False signals induced by single-cell imputation.
复制标题

DOI:
10.12688/f1000research.16613.1
复制
发表时间:
2018-01-01
期刊:
影响因子:
--
通讯作者:
Hemberg, Martin
Hemberg, Martin
中科院分区:
其他
文献类型:
--
作者:
Andrews, Tallulah S;Hemberg, Martin

文献摘要

被引文献

相似文献

背景:单细胞RNASeq是以单个细胞的分辨率测量基因表达的强大工具。分析这些数据的一个重大挑战是大量的零值,代表缺失数据或没有表达。已经提出了几种估算方法来处理这个问题,但由于这些方法通常依赖于所考虑的数据集的固有结构,它们可能无法提供任何额外的信息。方法:我们评估了用五种不同方法插补数据时产生假阳性或不可重现结果的风险。我们将每种方法应用于各种模拟数据集以及排列的真实的单细胞RNASeq数据集,并考虑假阳性基因-基因相关性和差异表达基因的数量。使用Tabula Muris数据库中匹配的10 X Chromium和Smartseq 2数据,我们检查了插补前后标记物的重现性。结果:由插补引入的假阳性信号的程度因方法而异。基于数据平滑的方法,MAGIC和knn-smooth,在真实的和模拟数据中都产生了非常高的误报率。基于模型的插补方法通常产生较少的假阳性,但这取决于数据集符合基础模型的程度。此外,在匹配数据中,只有SAVER显示出与未插补数据相当的重现性。结论:单细胞RNASeq数据的插补引入了可能产生假阳性结果的循环。因此,应用于插补数据的统计检验应谨慎对待。按效应大小进行的附加过滤可以减少但不能完全消除这些效应。在我们考虑的方法中,SAVER产生错误或不可重现结果的可能性最小,因此如果有必要进行插补,则应优先于其他方法。
Background: Single-cell RNASeq is a powerful tool for measuring gene expression at the resolution of individual cells. A significant challenge in the analysis of this data is the large amount of zero values, representing either missing data or no expression. Several imputation approaches have been proposed to deal with this issue, but since these methods generally rely on structure inherent to the dataset under consideration they may not provide any additional information. Methods: We evaluated the risk of generating false positive or irreproducible results when imputing data with five different methods. We applied each method to a variety of simulated datasets as well as to permuted real single-cell RNASeq datasets and consider the number of false positive gene-gene correlations and differentially expressed genes. Using matched 10X Chromium and Smartseq2 data from the Tabula Muris database we examined the reproducibility of markers before and after imputation. Results: The extent of false-positive signals introduced by imputation varied considerably by method. Data smoothing based methods, MAGIC and knn-smooth, generated a very high number of false-positives in both real and simulated data. Model-based imputation methods typically generated fewer false-positives but this varied greatly depending on how well datasets conformed to the underlying model. Furthermore, only SAVER exhibited reproducibility comparable to unimputed data across matched data. Conclusions: Imputation of single-cell RNASeq data introduces circularity that can generate false-positive results. Thus, statistical tests applied to imputed data should be treated with care. Additional filtering by effect size can reduce but not fully eliminate these effects. Of the methods we considered, SAVER was the least likely to generate false or irreproducible results, thus should be favoured over alternatives if imputation is necessary.