Inference of species phylogenies from bi-allelic markers using pseudo-likelihood.

Inference of species phylogenies from bi-allelic markers using pseudo-likelihood.
复制标题

DOI:
10.1093/bioinformatics/bty295
复制
发表时间:
2018-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Nakhleh L
Nakhleh L
中科院分区:
其他
文献类型:
--
作者:
Zhu J;Nakhleh L

文献摘要

参考文献

被引文献

相似文献

系统发生网络代表了网状的进化历史。统计方法推断下,他们的多物种合并最近已经开发出来。一种特别强大的方法使用由双等位基因标记(例如单核苷酸多态性数据)组成的数据,并允许系统发生网络的精确似然计算,同时对每个标记的所有可能的基因树进行数值积分。虽然该方法在估计网络及其参数方面具有良好的精度,但似然计算仍然是一个主要的计算瓶颈,并限制了该方法的适用性。在本文中,我们首先演示了为什么网络的似然计算与树相比需要多几个数量级的时间。然后,我们提出了一种基于伪似然使用双等位基因标记的系统发育网络的推断方法。我们通过模拟数据上的伪似然计算证明了系统发育网络推理的可扩展性和准确性。此外,我们证明了方面的鲁棒性的方法违反所采用的统计模型的基本假设。最后,我们展示了该方法的生物数据的应用。所提出的方法允许分析较大的数据集的类群和网状事件的数量。虽然伪似然之前已经提出了由基因树组成的数据,但这里的工作直接使用序列数据,提供了我们讨论的几个优点。这些方法已在PhyloNet(http://bioinfocs.rice.edu/phylonet)中实施。
Phylogenetic networks represent reticulate evolutionary histories. Statistical methods for their inference under the multispecies coalescent have recently been developed. A particularly powerful approach uses data that consist of bi-allelic markers (e.g. single nucleotide polymorphism data) and allows for exact likelihood computations of phylogenetic networks while numerically integrating over all possible gene trees per marker. While the approach has good accuracy in terms of estimating the network and its parameters, likelihood computations remain a major computational bottleneck and limit the method’s applicability. In this article, we first demonstrate why likelihood computations of networks take orders of magnitude more time when compared to trees. We then propose an approach for inference of phylogenetic networks based on pseudo-likelihood using bi-allelic markers. We demonstrate the scalability and accuracy of phylogenetic network inference via pseudo-likelihood computations on simulated data. Furthermore, we demonstrate aspects of robustness of the method to violations in the underlying assumptions of the employed statistical model. Finally, we demonstrate the application of the method to biological data. The proposed method allows for analyzing larger datasets in terms of the numbers of taxa and reticulation events. While pseudo-likelihood had been proposed before for data consisting of gene trees, the work here uses sequence data directly, offering several advantages as we discuss. The methods have been implemented in PhyloNet (http://bioinfocs.rice.edu/phylonet).
DOI: 10.1126/science.1258524
发表时间: 2015-01-02
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Fontaine MC;Pease JB;Steele A;Waterhouse RM;Neafsey DE;Sharakhov IV;Jiang X;Hall AB;Catteruccia F;Kakani E;Mitchell SN;Wu YC;Smith HA;Love RR;Lawniczak MK;Slotman MA;Emrich SJ;Hahn MW;Besansky NJ
通讯作者: Besansky NJ
DOI: 10.1038/nrg3936
发表时间: 2015-06
期刊: Nature reviews. Genetics
影响因子: --
作者:
通讯作者: --
DOI: 10.1093/molbev/mss086
发表时间: 2012-08-01
影响因子: 10.7
作者:
Bryant, David;Bouckaert, Remco;RoyChoudhury, Arindam
通讯作者: RoyChoudhury, Arindam
DOI: 10.1007/bf01734359
发表时间: 1981-01-01
影响因子: 3.9
作者:
FELSENSTEIN, J
通讯作者: FELSENSTEIN, J
DOI: 10.1002/bies.201500149
发表时间: 2016-02
期刊: BioEssays : news and reviews in molecular, cellular and developmental biology
影响因子: --
作者:
Mallet J;Besansky N;Hahn MW
通讯作者: Hahn MW