eFindSite: Improved prediction of ligand binding sites in protein models using meta-threading, machine learning and auxiliary ligands

eFindSite: Improved prediction of ligand binding sites in protein models using meta-threading, machine learning and auxiliary ligands
复制标题

DOI:
10.1007/s10822-013-9663-5
复制
发表时间:
2013-06-01
影响因子:
3.5
通讯作者:
Feinstein, Wei P.
Feinstein, Wei P.
中科院分区:
生物学3区
文献类型:
--
作者:
Brylinski, Michal;Feinstein, Wei P.

文献摘要

被引文献

相似文献

不同物种中大多数蛋白质的分子结构和功能尚未确定。这些基因产物的功能注释常常得益于蛋白质-配体相互作用的知识。为了实现这一目标,我们开发了eFindSite,一个改进版本的FINDSITE,旨在更有效地识别配体结合位点和残基,仅使用弱同源模板。它采用了一系列有效的算法,包括高度敏感的元线程方法,改进的聚类技术,先进的机器学习方法和可靠的置信度估计系统。根据目标蛋白质结构的质量,eFindSite在结合位点检测方面优于几何口袋检测算法15 - 40%,在结合残基预测方面优于几何口袋检测算法5 - 35%。此外,与FINDSITE相比,它在最困难的情况下多识别14%的结合残基。当识别出多个推定的结合口袋时,排序准确度为75 - 78%,通过包括从生物医学文献中提取的关于结合配体的辅助信息,其可以进一步提高3 - 4%。作为第一个跨基因组的应用,我们描述了整个蛋白质组的大肠杆菌的结构建模和结合位点预测。仔细校准的置信度估计强烈表明,高度可靠的配体结合预测的大多数基因产物,因此eFindSite持有大规模基因组注释和药物开发项目的重大承诺。eFindSite可在www.example.com上免费向学术界提供。
Molecular structures and functions of the majority of proteins across different species are yet to be identified. Much needed functional annotation of these gene products often benefits from the knowledge of protein-ligand interactions. Towards this goal, we developed eFindSite, an improved version of FINDSITE, designed to more efficiently identify ligand binding sites and residues using only weakly homologous templates. It employs a collection of effective algorithms, including highly sensitive meta-threading approaches, improved clustering techniques, advanced machine learning methods and reliable confidence estimation systems. Depending on the quality of target protein structures, eFindSite outperforms geometric pocket detection algorithms by 15-40 % in binding site detection and by 5-35 % in binding residue prediction. Moreover, compared to FINDSITE, it identifies 14 % more binding residues in the most difficult cases. When multiple putative binding pockets are identified, the ranking accuracy is 75-78 %, which can be further improved by 3-4 % by including auxiliary information on binding ligands extracted from biomedical literature. As a first across-genome application, we describe structure modeling and binding site prediction for the entire proteome of Escherichia coli. Carefully calibrated confidence estimates strongly indicate that highly reliable ligand binding predictions are made for the majority of gene products, thus eFindSite holds a significant promise for large-scale genome annotation and drug development projects. eFindSite is freely available to the academic community at http://www.brylinski.org/efindsite.