Pandit: a database of protein and associated nucleotide domains with inferred trees

Pandit: a database of protein and associated nucleotide domains with inferred trees
复制标题

DOI:
10.1093/bioinformatics/btg188
复制
发表时间:
2003-08-12
期刊:
影响因子:
5.8
通讯作者:
Goldman, N
Goldman, N
中科院分区:
生物学3区
文献类型:
--
作者:
Whelan, S;de Bakker, PIW;Goldman, N

文献摘要

被引文献

相似文献

动机:一个大的,高质量的同源序列比对数据库,以及它们相应的系统发育树的良好估计,将是那些研究遗传学的宝贵资源。它将使研究人员能够在各种各样的序列中比较当前和新的序列进化模型。大量的数据可以为研究序列进化的新模型和方法提供灵感,并且可以允许关于不同分子过程对进化的相对影响的一般性陈述。结果:Pandit 7.6数据库包含4341个家族的序列,这些序列来自同源蛋白质结构域家族的氨基酸比对的Pfam数据库的种子比对(Bateman等人,2002年)。Pandit中的每个家族包括与相应Pfam家族种子比对匹配的氨基酸序列比对、当可以回收时含有Pfam比对的编码序列的DNA序列比对(总体上,82.9%的序列取自Pfam)以及仅限于可以回收DNA序列的那些序列的氨基酸序列比对。每个比对都有一个与之相关的系统发育树的估计。树的拓扑结构是使用基于进化距离的最大似然估计的邻居连接方法获得的,然后使用标准的最大似然方法计算分支长度。
Motivation: A large, high-quality database of homologous sequence alignments with good estimates of their corresponding phylogenetic trees will be a valuable resource to those studying phylogenetics. It will allow researchers to compare current and new models of sequence evolution across a large variety of sequences. The large quantity of data may provide inspiration for new models and methodology to study sequence evolution and may allow general statements about the relative effect of different molecular processes on evolution.Results: The Pandit 7.6 database contains 4341 families of sequences derived from the seed alignments of the Pfam database of amino acid alignments of families of homologous protein domains (Bateman et al., 2002). Each family in Pandit includes an alignment of amino acid sequences that matches the corresponding Pfam family seed alignment, an alignment of DNA sequences that contain the coding sequence of the Pfam alignment when they can be recovered (overall, 82.9% of sequences taken from Pfam) and the alignment of amino acid sequences restricted to only those sequences for which a DNA sequence could be recovered. Each of the alignments has an estimate of the phylogenetic tree associated with it. The tree topologies were obtained using the neighbor joining method based on maximum likelihood estimates of the evolutionary distances, with branch lengths then calculated using a standard maximum likelihood approach.