A novel statistical ligand-binding site predictor: application to ATP-binding sites

A novel statistical ligand-binding site predictor: application to ATP-binding sites
复制标题

DOI:
10.1093/protein/gzi006
复制
发表时间:
2005-02-01
影响因子:
2.4
通讯作者:
Sun, ZR
Sun, ZR
中科院分区:
生物学4区
文献类型:
--
作者:
Guo, T;Shi, YX;Sun, ZR

文献摘要

被引文献

相似文献

结构基因组学计划正在导致新确定的蛋白质3D结构的快速增长,其功能表征可能仍然不充分。为了深入了解结构可用的新兴蛋白质的可能作用和/或补充生化研究,已经开发了各种计算方法,用于筛选和预测原始结构数据中的配体结合位点,包括统计模式分类技术。在本文中,我们报告了一种新的蛋白质配体结合位点的统计描述符(定向壳模型),它利用了紧邻结合位点中心的各种结构和物理化学特征的距离和角位置分布。使用支持向量机(SVM)作为分类器,我们的模型在全蛋白扫描测试中识别出69%的atp结合位点,在真核蛋白中准确率特别高。我们提出这种特征提取和机器学习过程可以筛选出具有配体结合能力的候选蛋白质,并可以为单个蛋白质提供有价值的生化信息。
Structural genomics initiatives are leading to rapid growth in newly determined protein 3D structures, the functional characterization of which may still be inadequate. As an attempt to provide insights into the possible roles of the emerging proteins whose structures are available and/or to complement biochemical research, a variety of computational methods have been developed for the screening and prediction of ligand-binding sites in raw structural data, including statistical pattern classification techniques. In this paper, we report a novel statistical descriptor (the Oriented Shell Model) for protein ligand-binding sites, which utilizes the distance and angular position distribution of various structural and physicochemical features present in immediate proximity to the center of a binding site. Using the support vector machine (SVM) as the classifier, our model identified 69% of the ATP-binding sites in whole-protein scanning tests and in eukaryotic proteins the accuracy is particularly high. We propose that this feature extraction and machine learning procedure can screen out ligand-binding-capable protein candidates and can yield valuable biochemical information for individual proteins.