Prospects for Building Large Timetrees Using Molecular Data with Incomplete Gene Coverage among Species

Prospects for Building Large Timetrees Using Molecular Data with Incomplete Gene Coverage among Species
复制标题

DOI:
10.1093/molbev/msu200
复制
发表时间:
2014-09-01
影响因子:
10.7
通讯作者:
Kumar, Sudhir
Kumar, Sudhir
中科院分区:
生物学1区
文献类型:
--
作者:
Filipski, Alan;Murillo, Oscar;Kumar, Sudhir

文献摘要

被引文献

相似文献

科学家们正在从越来越多的物种和基因中收集序列数据集,以建立全面的时间树。然而,一些物种和基因组合的数据往往是不可用的,缺失数据的比例往往是很大的数据集包含许多基因和物种。令人惊讶的是,目前还没有一个系统的分析的稀疏程度的物种genematrix上的分歧时间估计的准确性的影响。在这里,我们提出了计算机模拟和经验数据分析的结果,以量化缺失的基因数据对大型基因组中发散时间估计的影响。我们发现,即使在大多数物种的大多数基因序列缺失的情况下,对分歧时间的估计也是稳健的。通过对这些极其稀疏的数据集的分析,我们发现,最严重的错误发生在树中的节点上,这些节点在所讨论的节点的直接后代分支中的任何一对物种中都没有共同的基因。这些有问题的节点可以很容易地检测之前,计算分析的基础上,只有输入序列比对和树的拓扑结构。我们的结论是,最好使用更大的比对,因为添加基因和物种的比对增加了基因的数量,可用于估计分歧事件在树的深处,并提高他们的时间估计。
Scientists are assembling sequence data sets from increasing numbers of species and genes to build comprehensive timetrees. However, data are often unavailable for some species and gene combinations, and the proportion of missing data is often large for data sets containing many genes and species. Surprisingly, there has not been a systematic analysis of the effect of the degree of sparseness of the species-genematrix on the accuracy of divergence time estimates. Here, we present results from computer simulations and empirical data analyses to quantify the impact of missing gene data on divergence time estimation in large phylogenies. We found that estimates of divergence times were robust even when sequences from a majority of genes for most of the species were absent. From the analysis of such extremely sparse data sets, we found that the most egregious errors occurred for nodes in the tree that had no common genes for any pair of species in the immediate descendant clades of the node in question. These problematic nodes can be easily detected prior to computational analyses based only on the input sequence alignment and the tree topology. We conclude that it is best to use larger alignments, because adding both genes and species to the alignment augments the number of genes available for estimating divergence events deep in the tree and improves their time estimates.