Approximate likelihood-ratio test for branches: A fast, accurate, and powerful alternative

Approximate likelihood-ratio test for branches: A fast, accurate, and powerful alternative
复制标题

DOI:
10.1080/10635150600755453
复制
发表时间:
2006-08-01
期刊:
影响因子:
6.5
通讯作者:
Gascuel, Olivier
Gascuel, Olivier
中科院分区:
生物学1区
文献类型:
--
作者:
Anisimova, Maria;Gascuel, Olivier

文献摘要

被引文献

相似文献

我们回顾了基于分子数据重建的进化树分支的统计测试。本文提出了一种新的、快速的分支近似似然比检验(ALRT),作为分支支持度的非参数Bootstrap和贝叶斯估计的竞争性替代。ALRT基于传统LRT的思想,零假设对应于推断的分支具有长度0的假设。我们证明了LRT统计量是由1/2chi(2)(0)+1/2chi(2)(1)分布所得到的最大三个随机变量的渐近分布。内部分支的新ALRT使用这种分布进行显著性检验,但检验统计量以一种稍微保守但实用的方式近似为2(L(1)-L(2)),即最好的树对应的最大对数似然值与感兴趣的分支周围的次优拓扑排列之间的差的两倍。这样的测试是快速的,因为对数似然值L(2)是通过仅在感兴趣的分支和四个相邻分支上进行优化来计算的,而其他参数被固定在与最佳ML树相对应的它们的最佳值。在具有不同长度序列的模拟4分类单元、12分类单元和100分类单元的数据集上研究了新测试的性能。ALRT被证明是准确的、强大的,并且对于某些违反模型假设的情况是健壮的。ALRT是在最近的快速最大似然树估计程序PHYML(Guindon和Gascuel,2003)使用的算法内实现的。[准确性;分支支持;似然比检验;系统发育重建;能力。]
We revisit statistical tests for branches of evolutionary trees reconstructed upon molecular data. A new, fast, approximate likelihood-ratio test (aLRT) for branches is presented here as a competitive alternative to nonparametric bootstrap and Bayesian estimation of branch support. The aLRT is based on the idea of the conventional LRT, with the null hypothesis corresponding to the assumption that the inferred branch has length 0. We show that the LRT statistic is asymptotically distributed as a maximum of three random variables drawn from the 1/2 chi(2)(0) + 1/2 chi(2)(1) distribution. The new aLRT of interior branch uses this distribution for significance testing, but the test statistic is approximated in a slightly conservative but practical way as 2(l(1)-l(2)), i.e., double the difference between the maximum log-likelihood values corresponding to the best tree and the second best topological arrangement around the branch of interest. Such a test is fast because the log-likelihood value l(2) is computed by optimizing only over the branch of interest and the four adjacent branches, whereas other parameters are fixed at their optimal values corresponding to the best ML tree. The performance of the new test was studied on simulated 4-, 12-, and 100-taxon data sets with sequences of different lengths. The aLRT is shown to be accurate, powerful, and robust to certain violations of model assumptions. The aLRT is implemented within the algorithm used by the recent fast maximum likelihood tree estimation program PHYML (Guindon and Gascuel, 2003). [Accuracy; branch support; likelihood-ratio test; phylogeny reconstruction; power.]