Evidence of Statistical Inconsistency of Phylogenetic Methods in the Presence of Multiple Sequence Alignment Uncertainty.

Evidence of Statistical Inconsistency of Phylogenetic Methods in the Presence of Multiple Sequence Alignment Uncertainty.
复制标题

DOI:
10.1093/gbe/evv127
复制
发表时间:
2015-07-01
影响因子:
3.3
通讯作者:
Whelan S
Whelan S
中科院分区:
生物学2区
文献类型:
--
作者:
Md Mukarram Hossain AS;Blackburne BP;Shah A;Whelan S

文献摘要

被引文献

相似文献

进化研究通常使用两个步骤来研究序列数据。第一步估计多序列比对(MSA),第二步应用系统发育方法来询问MSA的进化问题。现代系统发育方法使用最大似然或贝叶斯推断来推断进化参数,由描述树上序列变化的概率替代模型介导。这些方法的统计特性意味着更多的数据直接转化为下游结果的置信度增加,前提是替代模型足够并且MSA正确。许多研究已经调查了系统发育方法的鲁棒性存在的替代模型误指定,但很少有研究这些方法的统计特性时,MSA是未知的。本模拟研究探讨了完整的两个步骤的过程中推断序列的分歧和系统发育树拓扑结构的统计特性。核苷酸和氨基酸分析都受到比对步骤的负面影响,既通过不准确的指导树估计,也通过过度拟合该指导树。对于许多比对工具,当将额外的序列添加到分析中时,这些影响变得更加明显。核苷酸序列是特别敏感的,MSA错误导致统计支持长分支吸引力的文物,这通常是与总取代模型误指定。氨基酸MSA更稳健,但倾向于任意解析多分叉,而有利于指导树。没有推理策略产生一致准确的估计序列之间的分歧,虽然氨基酸的MSA再次比它们的核苷酸对应物更准确。最后,我们就如何限制MSA不确定性对进化推理的影响提出了一些切实可行的建议。
Evolutionary studies usually use a two-step process to investigate sequence data. Step one estimates a multiple sequence alignment (MSA) and step two applies phylogenetic methods to ask evolutionary questions of that MSA. Modern phylogenetic methods infer evolutionary parameters using maximum likelihood or Bayesian inference, mediated by a probabilistic substitution model that describes sequence change over a tree. The statistical properties of these methods mean that more data directly translates to an increased confidence in downstream results, providing the substitution model is adequate and the MSA is correct. Many studies have investigated the robustness of phylogenetic methods in the presence of substitution model misspecification, but few have examined the statistical properties of those methods when the MSA is unknown. This simulation study examines the statistical properties of the complete two-step process when inferring sequence divergence and the phylogenetic tree topology. Both nucleotide and amino acid analyses are negatively affected by the alignment step, both through inaccurate guide tree estimates and through overfitting to that guide tree. For many alignment tools these effects become more pronounced when additional sequences are added to the analysis. Nucleotide sequences are particularly susceptible, with MSA errors leading to statistical support for long-branch attraction artifacts, which are usually associated with gross substitution model misspecification. Amino acid MSAs are more robust, but do tend to arbitrarily resolve multifurcations in favor of the guide tree. No inference strategies produce consistently accurate estimates of divergence between sequences, although amino acid MSAs are again more accurate than their nucleotide counterparts. We conclude with some practical suggestions about how to limit the effect of MSA uncertainty on evolutionary inference.