Protein subfamily assignment using the Conserved Domain Database.

Protein subfamily assignment using the Conserved Domain Database.
复制标题

DOI:
10.1186/1756-0500-1-114
复制
发表时间:
2008-11-14
期刊:
影响因子:
1.8
通讯作者:
Marchler-Bauer, Aron
Marchler-Bauer, Aron
中科院分区:
其他
文献类型:
--
作者:
Fong, Jessica H;Marchler-Bauer, Aron

文献摘要

被引文献

相似文献

结构域是蛋白质进化上的保守单位,广泛用于蛋白质序列的分类和蛋白质功能的推断。通常,两个或更多个重叠的结构域模型匹配蛋白质序列的区域。因此,需要程序来为蛋白质选择适当的结构域注释。在这里,我们提出了一种方法,用于分配NCBI策划域从策划域数据库(CDD),考虑到组织的域到层次结构的同源域模型。我们对NCBI策划的领域分配的比对分数的分析表明,在密切相关的模型中识别正确的模型比在不重叠的领域模型之间进行选择更困难。我们发现,基于排序分数和特定领域的阈值的简单的分类是有效的,在减少分类错误。事实上,在我们的测试集中,由于缺失的域子家族被更通用的域分配所取代,从而消除了数据库中大量的错误,因此,在当前的错误分类中,几乎有90%的错误分类是由分类法导致的。我们提出的结构域亚家族分配规则已被纳入CD-Search软件分配CDD结构域查询蛋白质序列,并显着提高了预先计算的结构域注释的蛋白质序列在NCBI的NICUZ资源。
Domains, evolutionarily conserved units of proteins, are widely used to classify protein sequences and infer protein function. Often, two or more overlapping domain models match a region of a protein sequence. Therefore, procedures are required to choose appropriate domain annotations for the protein. Here, we propose a method for assigning NCBI-curated domains from the Curated Domain Database (CDD) that takes into account the organization of the domains into hierarchies of homologous domain models. Our analysis of alignment scores from NCBI-curated domain assignments suggests that identifying the correct model among closely related models is more difficult than choosing between non-overlapping domain models. We find that simple heuristics based on sorting scores and domain-specific thresholds are effective at reducing classification error. In fact, in our test set, the heuristics result in almost 90% of current misclassifications due to missing domain subfamilies being replaced by more generic domain assignments, thereby eliminating a significant amount of error within the database. Our proposed domain subfamily assignment rule has been incorporated into the CD-Search software for assigning CDD domains to query protein sequences and has significantly improved pre-calculated domain annotations on protein sequences in NCBI's Entrez resource.