Prospects for building the tree of life from large sequence databases

Prospects for building the tree of life from large sequence databases
复制标题

DOI:
10.1126/science.1102036
复制
发表时间:
2004-11-12
期刊:
影响因子:
56.9
通讯作者:
Sanderson, MJ
Sanderson, MJ
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Driskell, AC;Ané, C;Sanderson, MJ

文献摘要

被引文献

相似文献

我们评估了从 Swiss-Prot 和 GenBank 采样的类似 300,000 个蛋白质序列的系统发育潜力。尽管这些数据中只有一小部分具有潜在的系统发育信息,但该子集保留了原始分类多样性的很大一部分。数据库中的采样偏差需要构建包含大量缺失条目的系统发育数据集。然而,对两个“超级矩阵”的分析表明,即使数据集缺失的数据高达 92%,也可以提供对生命树的广泛部分的洞察。
We assess the phylogenetic potential of similar to300,000 protein sequences sampled from Swiss-Prot and GenBank. Although only a small subset of these data was potentially phylogenetically informative, this subset retained a substantial fraction of the original taxonomic diversity. Sampling biases in the databases necessitate building phylogenetic data sets that have large numbers of missing entries. However, an analysis of two "supermatrices" suggests that even data sets with as much as 92% missing data can provide insights into broad sections of the tree of life.