On single and multiple models of protein families for the detection of remote sequence relationships.

On single and multiple models of protein families for the detection of remote sequence relationships.
复制标题

DOI:
10.1186/1471-2105-7-48
复制
发表时间:
2006-01-31
期刊:
影响因子:
3
通讯作者:
Saqi MA
Saqi MA
中科院分区:
生物学4区
文献类型:
--
作者:
Casbon JA;Saqi MA

文献摘要

参考文献

被引文献

相似文献

检测未知功能的蛋白质序列与其功能已被表征的序列之间的关系使得能够转移功能注释。然而,在许多情况下,这些关系不能从两个序列的直接比较中容易地确定。比较序列谱的方法已经显示出改善这些远程序列关系的检测。然而,建立一组已知序列图谱的最佳方法尚未建立。在这里,我们研究如何建立的配置文件的类型影响其性能,无论是在检测远程同源物,并在由此产生的对齐精度。特别是,我们考虑是否更好地使用一个单一的基于结构的比对,代表所有已知的情况下的超家族,或使用多个基于序列的配置文件,每个代表一个单独的成员的超家族的蛋白质超家族模型。使用配置文件的远程同源检测方法,我们基准性能的单结构为基础的超家族模型和多域模型。平均而言,在所有超家族中,使用截断接收算子特征(ROC 5),我们发现多域模型优于单超家族模型,除了在低错误率下,两种模型的行为方式相似。然而,有一个广泛的性能取决于超家族。对于12%的所有超家族的ROC5值的超家族模型是大于0.2以上的域模型和10%的超家族的域模型表现出类似的改善性能的超家族模型。使用一个敏感的配置文件配置文件的方法,我们已经研究了性能的单一结构为基础的模型和多个序列模型(域模型)在检测远程超家族成员。我们发现,总体而言,多个模型在识别中表现更好,尽管基于单一结构的模型显示出更好的对齐精度。
The detection of relationships between a protein sequence of unknown function and a sequence whose function has been characterised enables the transfer of functional annotation. However in many cases these relationships can not be identified easily from direct comparison of the two sequences. Methods which compare sequence profiles have been shown to improve the detection of these remote sequence relationships. However, the best method for building a profile of a known set of sequences has not been established. Here we examine how the type of profile built affects its performance, both in detecting remote homologs and in the resulting alignment accuracy. In particular, we consider whether it is better to model a protein superfamily using a single structure-based alignment that is representative of all known cases of the superfamily, or to use multiple sequence-based profiles each representing an individual member of the superfamily. Using profile-profile methods for remote homolog detection we benchmark the performance of single structure-based superfamily models and multiple domain models. On average, over all superfamilies, using a truncated receiver operator characteristic (ROC5) we find that multiple domain models outperform single superfamily models, except at low error rates where the two models behave in a similar way. However there is a wide range of performance depending on the superfamily. For 12% of all superfamilies the ROC5 value for superfamily models is greater than 0.2 above the domain models and for 10% of superfamilies the domain models show a similar improvement in performance over the superfamily models. Using a sensitive profile-profile method we have investigated the performance of single structure-based models and multiple sequence models (domain models) in detecting remote superfamily members. We find that overall, multiple models perform better in recognition although single structure-based models display better alignment accuracy.
DOI: 10.1016/j.jmb.2003.10.025
发表时间: 2003-12-12
影响因子: 5.6
作者:
Tang, CL;Xie, L;Honig, B
通讯作者: Honig, B
DOI: 10.1093/bib/3.3.296
发表时间: 2002-09-01
影响因子: 9.5
作者:
Mangalam, Harry
通讯作者: Mangalam, Harry
DOI: 10.1093/bioinformatics/bti125
发表时间: 2005-04-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Söding, J
通讯作者: Söding, J
DOI: 10.1006/jmbi.2001.5080
发表时间: 2001-11-02
影响因子: 5.6
作者:
Gough, J;Karplus, K;Chothia, C
通讯作者: Chothia, C
DOI: 10.1110/ps.03197403
发表时间: 2003-10
期刊: Protein science : a publication of the Protein Society
影响因子: --
作者:
Sadreyev RI;Baker D;Grishin NV
通讯作者: Grishin NV