Protein function prediction using domain families.

Protein function prediction using domain families.
复制标题

DOI:
10.1186/1471-2105-14-s3-s5
复制
发表时间:
2013
期刊:
影响因子:
3
通讯作者:
Orengo CA
Orengo CA
中科院分区:
生物学4区
文献类型:
--
作者:
Rentzsch R;Orengo CA

文献摘要

被引文献

相似文献

在这里,我们评估了使用结构域家族来预测整个蛋白质的功能。这些‘功能家族’(FunFam)是通过在后一步中依赖可用的高质量基因本体论(GO)注释数据,使用将序列聚类与监督聚类评估相结合的协议来获得的。本质上,该协议根据其亲本蛋白质的GO注释将属于同一超家族的结构域序列分组为家族。基于酶序列的初步测试证实,FunFams与酶(结构域)家族的相似性比单独通过序列聚类产生的家族要好得多。对于2011年的CAFA实验,我们进一步将FunFam与围棋术语进行了概率关联。所有的靶蛋白首先被提交给结构域超家族分配,然后是FunFam分配,最后是功能分配。后者包括多结构域靶蛋白的整合步骤。CAFA的结果使我们的基于领域的方法跻身31个竞争小组和56个预测方法的前十名,证实了它的性能优于简单的成对全蛋白序列比较。
Here we assessed the use of domain families for predicting the functions of whole proteins. These 'functional families' (FunFams) were derived using a protocol that combines sequence clustering with supervised cluster evaluation, relying on available high-quality Gene Ontology (GO) annotation data in the latter step. In essence, the protocol groups domain sequences belonging to the same superfamily into families based on the GO annotations of their parent proteins. An initial test based on enzyme sequences confirmed that the FunFams resemble enzyme (domain) families much better than do families produced by sequence clustering alone. For the CAFA 2011 experiment, we further associated the FunFams with GO terms probabilistically. All target proteins were first submitted to domain superfamily assignment, followed by FunFam assignment and, eventually, function assignment. The latter included an integration step for multi-domain target proteins. The CAFA results put our domain-based approach among the top ten of 31 competing groups and 56 prediction methods, confirming that it outperforms simple pairwise whole-protein sequence comparisons.