Accurate and efficient cell lineage tree inference from noisy single cell data: the maximum likelihood perfect phylogeny approach

Accurate and efficient cell lineage tree inference from noisy single cell data: the maximum likelihood perfect phylogeny approach
复制标题

DOI:
10.1093/bioinformatics/btz676
复制
发表时间:
2020-02-01
期刊:
影响因子:
5.8
通讯作者:
Wu, Yufeng
Wu, Yufeng
中科院分区:
生物学3区
文献类型:
--
作者:
Wu, Yufeng

文献摘要

被引文献

相似文献

动机:生物体中的细胞具有共同的进化历史,称为细胞谱系树。细胞谱系树可以从基因组变异位点的单细胞基因型推断出来。从嘈杂的单细胞数据中推断细胞谱系树是一个具有挑战性的计算问题。大多数现有的细胞谱系树推断方法都假设基因型具有统一的不确定性。一个关键的缺失方面是,真实的单细胞数据通常在个体基因型中具有不均匀的不确定性。此外,现有的方法通常是基于采样的,对于大数据来说可能非常慢。结果:在本文中,我们提出了一种称为 ScisTree 的新方法,它推断细胞谱系树并从嘈杂的单细胞基因型数据中调用基因型。与大多数现有方法不同,ScisTree 使用个体基因型的基因型概率(可以通过现有的单细胞基因型调用者计算)。 ScisTree 假设无限站点模型。考虑到具有个体化概率的不确定基因型,ScisTree 实现了一种快速启发式方法来推断细胞谱系树并调用基因型,从而实现所谓的完美系统发育并最大化基因型的可能性。通过仿真,我们表明 ScisTree 在推断树的准确性方面表现良好,并且比现有方法高效得多。 ScisTree 的效率使得新的应用成为可能,包括所谓的双峰插补。
Motivation: Cells in an organism share a common evolutionary history, called cell lineage tree. Cell lineage tree can be inferred from single cell genotypes at genomic variation sites. Cell lineage tree inference from noisy single cell data is a challenging computational problem. Most existing methods for cell lineage tree inference assume uniform uncertainty in genotypes. A key missing aspect is that real single cell data usually has non-uniform uncertainty in individual genotypes. Moreover, existing methods are often sampling based and can be very slow for large data.Results: In this article, we propose a new method called ScisTree, which infers cell lineage tree and calls genotypes from noisy single cell genotype data. Different from most existing approaches, ScisTree works with genotype probabilities of individual genotypes (which can be computed by existing single cell genotype callers). ScisTree assumes the infinite sites model. Given uncertain genotypes with individualized probabilities, ScisTree implements a fast heuristic for inferring cell lineage tree and calling the genotypes that allow the so-called perfect phylogeny and maximize the likelihood of the genotypes. Through simulation, we show that ScisTree performs well on the accuracy of inferred trees, and is much more efficient than existing methods. The efficiency of ScisTree enables new applications including imputation of the so-called doublets.