Protein-ligand binding site recognition using complementary binding-specific substructure comparison and sequence profile alignment

Protein-ligand binding site recognition using complementary binding-specific substructure comparison and sequence profile alignment
复制标题

DOI:
10.1093/bioinformatics/btt447
复制
发表时间:
2013-10-15
期刊:
影响因子:
5.8
通讯作者:
Zhang, Yang
Zhang, Yang
中科院分区:
生物学3区
文献类型:
--
作者:
Yang, Jianyi;Roy, Ambrish;Zhang, Yang

文献摘要

被引文献

相似文献

动机:蛋白质 - 配体结合位点的鉴定对于蛋白质功能注释和药物发现至关重要。然而,没有一种方法能够针对不同的蛋白质类型生成最佳的结合位点预测。互补预测的结合可能是解决该问题最可靠的方法。 结果:我们开发了两种新方法,一种基于结合特异性亚结构比较(TM - SITE),另一种基于序列谱比对(S - SITE),用于互补的结合位点预测。这些方法在一组包含814个天然、类药和金属离子分子的500个非冗余蛋白质上进行了测试。从低分辨率的蛋白质结构预测开始,这些方法成功识别出超过51%的结合残基,平均马修斯相关系数(MCC)显著高于(在学生t检验中P值>10⁻⁹)其他最先进的方法,包括COFACTOR、FINDSITE和ConCavity。当将TM - SITE和S - SITE与其他基于结构的程序相结合时,一种共识方法(COACH)可以比最佳的单一预测将MCC提高15%。COACH在最近的全社区COMEO实验中进行了检验,并在过去的22个单独数据集中始终被评为最佳方法,其曲线下面积得分比第二好的方法高22.5%。这些数据展示了一种新的强大的蛋白质 - 配体结合位点识别方法,该方法已准备好用于全基因组基于结构的功能注释。
Motivation: Identification of protein-ligand binding sites is critical to protein function annotation and drug discovery. However, there is no method that could generate optimal binding site prediction for different protein types. Combination of complementary predictions is probably the most reliable solution to the problem.Results: We develop two new methods, one based on binding-specific substructure comparison (TM-SITE) and another on sequence profile alignment (S-SITE), for complementary binding site predictions. The methods are tested on a set of 500 non-redundant proteins harboring 814 natural, drug-like and metal ion molecules. Starting from low-resolution protein structure predictions, the methods successfully recognize >51% of binding residues with average Matthews correlation coefficient (MCC) significantly higher (with P-value >10(-9) in student t-test) than other state-of-the-art methods, including COFACTOR, FINDSITE and ConCavity. When combining TM-SITE and S-SITE with other structure-based programs, a consensus approach (COACH) can increase MCC by 15% over the best individual predictions. COACH was examined in the recent community-wide COMEO experiment and consistently ranked as the best method in last 22 individual datasets with the Area Under the Curve score 22.5% higher than the second best method. These data demonstrate a new robust approach to protein-ligand binding site recognition, which is ready for genome-wide structure-based function annotations.