Characterization and prediction of residues determining protein functional specificity.

Characterization and prediction of residues determining protein functional specificity.
复制标题

DOI:
10.1093/bioinformatics/btn214
复制
发表时间:
2008-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Singh M
Singh M
中科院分区:
其他
文献类型:
--
作者:
Capra JA;Singh M

文献摘要

参考文献

被引文献

相似文献

动机:在一个同源蛋白质家族中,蛋白质可能被分成不同的亚型,这些亚型具有不同于整个家族的特定功能。通常,存在于少数序列位置的氨基酸决定了每种蛋白质特定的功能特异性。了解这些特异性决定位置(SDP)有助于蛋白质功能预测、药物设计和实验分析。已经引入了一些基于序列的计算方法来识别SDP;然而,由于已知的实验确定的SDP的数量有限,这些方法的进一步发展和评估受到了阻碍。结果:我们结合了几个生物信息学资源来自动化一个过程,通常是手动进行的,以建立SDP的数据集。由此产生的大型数据集,其中包括酶中的SDPs,使我们能够根据它们的物理化学和进化特性来表征SDPs。这也有利于基于序列的SDP预测方法的大规模评估。我们提出了一种简单的基于序列的SDP预测方法GroupSim,并令人惊讶地表明,它与现有的一组有代表性的方法具有竞争力。我们还描述了ConsWin,一种考虑了邻近氨基酸序列保守的启发式方法,并证明了它提高了在我们的酶SDP大数据集上测试的所有方法的性能。可获得性:数据集和GroupSim代码可在http://compbio.cs.princeton.edu/specificity/上在线获得。
Motivation: Within a homologous protein family, proteins may be grouped into subtypes that share specific functions that are not common to the entire family. Often, the amino acids present in a small number of sequence positions determine each protein's particular function-al specificity. Knowledge of these specificity determining positions (SDPs) aids in protein function prediction, drug design and experimental analysis. A number of sequence-based computational methods have been introduced for identifying SDPs; however, their further development and evaluation have been hindered by the limited number of known experimentally determined SDPs. Results: We combine several bioinformatics resources to automate a process, typically undertaken manually, to build a dataset of SDPs. The resulting large dataset, which consists of SDPs in enzymes, enables us to characterize SDPs in terms of their physicochemical and evolution-ary properties. It also facilitates the large-scale evaluation of sequence-based SDP prediction methods. We present a simple sequence-based SDP prediction method, GroupSim, and show that, surprisingly, it is competitive with a representative set of current methods. We also describe ConsWin, a heuristic that considers sequence conservation of neighboring amino acids, and demonstrate that it improves the performance of all methods tested on our large dataset of enzyme SDPs. Availability: Datasets and GroupSim code are available online at http://compbio.cs.princeton.edu/specificity/ Contact: msingh@cs.princeton.edu Supplementary information: Supplementary data are available at Bioinformatics online.
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1186/gb-2006-7-1-r8
发表时间: 2006
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Brown, Shoshana D;Gerlt, John A;Seffernick, Jennifer L;Babbitt, Patricia C
通讯作者: Babbitt, Patricia C
Pfam:氏族、网络工具和服务。
DOI: 10.1093/nar/gkj149
发表时间: 2006-01-01
影响因子: 14.9
作者:
Finn, Robert D.;Mistry, Jaina;Schuster-Bockler, Benjamin;Griffiths-Jones, Sam;Hollich, Volker;Lassmann, Timo;Moxon, Simon;Marshall, Mhairi;Khanna, Ajay;Durbin, Richard;Eddy, Sean R.;Sonnhammer, Erik L. L.;Bateman, Alex
通讯作者: Bateman, Alex
DOI: 10.1110/ps.03191704
发表时间: 2004-02-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
Kalinina, OV;Mironov, AA;Rakhmaninova, AB
通讯作者: Rakhmaninova, AB
DOI: 10.1186/1471-2105-6-284
发表时间: 2005-11-30
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Mayer, KM;McCorkle, SR;Shanklin, J
通讯作者: Shanklin, J