Inference of reticulate evolutionary histories by maximum likelihood: the performance of information criteria.

Inference of reticulate evolutionary histories by maximum likelihood: the performance of information criteria.
复制标题

DOI:
10.1186/1471-2105-13-s19-s12
复制
发表时间:
2012
期刊:
影响因子:
3
通讯作者:
Nakhleh L
Nakhleh L
中科院分区:
生物学4区
文献类型:
--
作者:
Park HJ;Nakhleh L

文献摘要

相似文献

三十多年来,极大似然被广泛用于从分子数据推断系统发育树。当网状进化事件发生时,几个基因组区域可能具有相互冲突的进化历史,系统发育网络可能为代表基因组或物种的进化历史提供更充分的模型。对于这种情况,提出了一个最大似然(ML)模型,并解释了基因组区域内的突变和区域间的网状结构。然而,该模型在推断网络演化信息和影响该性能的特性方面的性能尚未得到研究。本文研究了在机器学习条件下网状事件的演化直径和高度对其可识别性的影响,发现两者,尤其是直径对其可识别性有显著影响。此外,我们发现在网状边缘转移的基因数量(可以概括为“非重组基因组区域”的概念)会影响其可检测性。最后但并非最不重要的是,系统发育网络的一个基本挑战是,它们允许任意程度的复杂性,从而产生模型选择问题。为了解决这个问题,我们研究了赤池信息准则(AIC)和贝叶斯信息准则(BIC)这两个信息准则的性能。我们发现BIC在控制模型复杂性和防止ML严重高估网状事件数量方面表现良好。我们的研究结果表明,BIC为推断网状进化历史提供了一个很好的框架。然而,在解释推理的准确性时,特别是对于具有特定进化特征的数据集,结果需要谨慎。
Maximum likelihood has been widely used for over three decades to infer phylogenetic trees from molecular data. When reticulate evolutionary events occur, several genomic regions may have conflicting evolutionary histories, and a phylogenetic network may provide a more adequate model for representing the evolutionary history of the genomes or species. A maximum likelihood (ML) model has been proposed for this case and accounts for both mutation within a genomic region and reticulation across the regions. However, the performance of this model in terms of inferring information about reticulate evolution and properties that affect this performance have not been studied. In this paper, we study the effect of the evolutionary diameter and height of a reticulation event on its identifiability under ML. We find both of them, particularly the diameter, have a significant effect. Further, we find that the number of genes (which can be generalized to the concept of "non-recombining genomic regions") that are transferred across a reticulation edge affects its detectability. Last but not least, a fundamental challenge with phylogenetic networks is that they allow an arbitrary level of complexity, giving rise to the model selection problem. We investigate the performance of two information criteria, the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), for addressing this problem. We find that BIC performs well in general for controlling the model complexity and preventing ML from grossly overestimating the number of reticulation events. Our results demonstrate that BIC provides a good framework for inferring reticulate evolutionary histories. Nevertheless, the results call for caution when interpreting the accuracy of the inference particularly for data sets with particular evolutionary features.