Impact of Missing Data on Phylogenies Inferred from Empirical Phylogenomic Data Sets

Impact of Missing Data on Phylogenies Inferred from Empirical Phylogenomic Data Sets
复制标题

DOI:
10.1093/molbev/mss208
复制
发表时间:
2013-01-01
影响因子:
10.7
通讯作者:
Philippe, Herve
Philippe, Herve
中科院分区:
生物学1区
文献类型:
--
作者:
Roure, Beatrice;Baurain, Denis;Philippe, Herve

文献摘要

被引文献

相似文献

测序技术的进步使研究人员能够组装更大的超矩阵,用于基因组推断。然而,目前的基因组学研究往往依赖于不完整的数据集,有些数据集有80%或更多的数据缺失(或模糊)。虽然早期的模拟表明,当使用足够大的数据集时,缺失数据本身不会损害系统发育推断,但Lemmon等人(Lemmon AR,Brown JM,Stanger-Hall K,Lemmon EM. 2009.模糊数据对最大似然法和贝叶斯推断法估计系统发育的影响。58:130-145)。最近,在一项研究中,对这种共识提出了质疑,这项研究是基于对简约-无信息的不完整字符的介绍。在这项工作中,我们经验性地重新评估的问题,缺失的数据在基因组学,同时探索可能的相互作用与模型的序列进化。首先,我们注意到,简约无信息的不完整字符实际上是信息的概率框架。考虑到这一点,对莱蒙数据集的重新分析对他们的结果给出了非常不同的解释,并表明他们的一些结论可能是没有根据的。其次,我们研究了在一个完整的超矩阵(126个基因× 39个物种)能够解决动物关系的缺失数据的逐步引入的效果。这些分析表明,丢失的数据扰动系统发育推断略超出预期的分辨率下降。特别是,它们通过减少可有效用于检测多个替代的物种的数量来加剧系统误差。因此,大型稀疏超矩阵比较小但不太完整的数据集对系统发育伪影更敏感,这表明实验设计旨在收集适度数量(类似于50)的高度覆盖基因。我们的结果进一步证实,包括不完整但短分支类群(即,缓慢进化的物种或封闭的外群体)可以帮助避免人工制品,正如模拟所预测的那样。最后,似乎选择适当的序列进化模型(例如,位点异质的CAT模型代替位点同质的WAG模型)比降低缺失数据水平更有利于系统发育的准确性。
Progress in sequencing technology allows researchers to assemble ever-larger supermatrices for phylogenomic inference. However, current phylogenomic studies often rest on patchy data sets, with some having 80% missing (or ambiguous) data or more. Though early simulations had suggested that missing data per se do not harm phylogenetic inference when using sufficiently large data sets, Lemmon et al. (Lemmon AR, Brown JM, Stanger-Hall K, Lemmon EM. 2009. The effect of ambiguous data on phylogenetic estimates obtained by maximum likelihood and Bayesian inference. Syst Biol. 58:130-145.) have recently cast doubt on this consensus in a study based on the introduction of parsimony-uninformative incomplete characters. In this work, we empirically reassess the issue of missing data in phylogenomics while exploring possible interactions with the model of sequence evolution. First, we note that parsimony-uninformative incomplete characters are actually informative in a probabilistic framework. A reanalysis of Lemmon's data set with this in mind gives a very different interpretation of their results and shows that some of their conclusions may be unfounded. Second, we investigate the effect of the progressive introduction of missing data in a complete supermatrix (126 genes x 39 species) capable of resolving animal relationships. These analyses demonstrate that missing data perturb phylogenetic inference slightly beyond the expected decrease in resolving power. In particular, they exacerbate systematic errors by reducing the number of species effectively available for the detection of multiple substitutions. Consequently, large sparse supermatrices are more sensitive to phylogenetic artifacts than smaller but less incomplete data sets, which argue for experimental designs aimed at collecting a modest number (similar to 50) of highly covered genes. Our results further confirm that including incomplete yet short-branch taxa (i.e., slowly evolving species or close outgroups) can help to eschew artifacts, as predicted by simulations. Finally, it appears that selecting an adequate model of sequence evolution (e.g., the site-heterogeneous CAT model instead of the site-homogeneous WAG model) is more beneficial to phylogenetic accuracy than reducing the level of missing data.