Probabilistic modeling of bifurcations in single-cell gene expression data using a Bayesian mixture of factor analyzers.

Probabilistic modeling of bifurcations in single-cell gene expression data using a Bayesian mixture of factor analyzers.
复制标题

DOI:
10.12688/wellcomeopenres.11087.1
复制
发表时间:
2017-03-15
影响因子:
--
通讯作者:
Yau C
Yau C
中科院分区:
其他
文献类型:
--
作者:
Campbell KR;Yau C

文献摘要

被引文献

相似文献

在单细胞转录组学数据中对分支进行建模已成为一个日益热门的研究领域。已经提出了几种从这类数据中推断分支结构的方法,但都依赖于启发式的非概率推断。在此,我们提出了首个基于贝叶斯分层混合因子分析器的用于此类推断的生成式、完全概率模型。尽管我们的模型实施了完整的马尔可夫链蒙特卡罗采样,但在大型数据集上仍表现出有竞争力的性能,并且其独特的分层先验结构能够自动确定驱动分支过程的基因。我们还提出了一种类似经验贝叶斯的扩展,用于处理单细胞RNA - seq数据中高水平的零膨胀情况,并量化此类模型何时有用。我们将我们的模型应用于真实和模拟的单细胞基因表达数据,并将结果与现有的拟时间方法进行比较。最后,我们在实际生物信息学分析的背景下讨论了这种统一的概率方法的优点和缺点。
Modeling bifurcations in single-cell transcriptomics data has become an increasingly popular field of research. Several methods have been proposed to infer bifurcation structure from such data, but all rely on heuristic non-probabilistic inference. Here we propose the first generative, fully probabilistic model for such inference based on a Bayesian hierarchical mixture of factor analyzers. Our model exhibits competitive performance on large datasets despite implementing full Markov-Chain Monte Carlo sampling, and its unique hierarchical prior structure enables automatic determination of genes driving the bifurcation process. We additionally propose an Empirical-Bayes like extension that deals with the high levels of zero-inflation in single-cell RNA-seq data and quantify when such models are useful. We apply or model to both real and simulated single-cell gene expression data and compare the results to existing pseudotime methods. Finally, we discuss both the merits and weaknesses of such a unified, probabilistic approach in the context practical bioinformatics analyses.