Discovering rules for protein-ligand specificity using support vector inductive logic programming.

Discovering rules for protein-ligand specificity using support vector inductive logic programming.
复制标题

DOI:
10.1093/protein/gzp035
复制
发表时间:
2009-09
期刊:
Protein engineering, design & selection : PEDS
影响因子:
--
通讯作者:
Sternberg MJ
Sternberg MJ
中科院分区:
其他
文献类型:
--
作者:
Kelley LA;Shrimpton PJ;Muggleton SH;Sternberg MJ

文献摘要

参考文献

被引文献

相似文献

结构基因组学计划正在迅速产生大量的蛋白质结构。比较建模也能够为许多蛋白质序列产生精确的结构模型。然而,对于许多已知的结构,功能尚未确定,并且在许多建模任务中,准确的结构模型不一定告诉我们功能。因此,迫切需要从结构确定功能的高通量方法。折叠蛋白质中关键氨基酸的空间排列,在表面或隐藏在裂缝中,通常是其生物功能的决定因素。分子生物学的一个中心目标是了解这些亚结构或表面与生物功能之间的关系,从而进行功能预测和功能设计。我们提出了一种新的通用方法,用于发现赋予特定配体的特异性的结合口袋的功能。使用最近开发的机器学习技术,该技术将归纳逻辑编程的规则发现方法与支持向量机的统计学习能力相结合,我们能够区分,准确率(90%)和召回率(86%)在仅给定结合口袋的骨架的几何形状和组成而没有结合口袋的骨架的情况下,使用对接。此外,我们还学习了控制这种特异性的规则,这些规则可以用于蛋白质功能设计方案。发现的规则的分析表明,结合口袋的关键特征可能与配体中的构象自由度有关。这一表述具有足够的普遍性,可适用于任何歧视性的约束性问题。所有程序和数据集均可在http://www.sbg.bio.ic.ac.uk/svilp_ligand/上免费向非商业用户提供。
Structural genomics initiatives are rapidly generating vast numbers of protein structures. Comparative modelling is also capable of producing accurate structural models for many protein sequences. However, for many of the known structures, functions are not yet determined, and in many modelling tasks an accurate structural model does not necessarily tell us about function. Thus there is a pressing need for high-throughput methods for determining function from structure. The spatial arrangement of key amino acids in a folded protein, on the surface or buried in clefts, are often the determinants of its biological function. A central aim of molecular biology is to understand the relationship between such substructures or surfaces and biological function, leading both to function prediction and function design. We present a new general method for discovering the features of binding pockets that confer specificity for particular ligands. Using a recently developed machine-learning technique which couples the rule-discovery approach of Inductive Logic Programming with the statistical learning power of Support Vector Machines, we are able to discriminate, with high precision (90%) and recall (86%) between pockets that bind FAD and those that bind NAD on a large benchmark set given only the geometry and composition of the backbone of the binding pocket without the use of docking. In addition we learn rules governing this specificity which can feed into protein functional design protocols. An analysis of the rules found suggest that key features of the binding pocket may be tied to conformational freedom in the ligand. The representation is sufficiently general to be applicable to any discriminatory binding problem. All programs and datasets are freely available to non-commercial users at http://www.sbg.bio.ic.ac.uk/svilp_ligand/.
DOI: 10.1073/pnas.93.1.438
发表时间: 1996-01-09
影响因子: 11.1
作者:
King, RD;Muggleton, SH;Sternberg, MJE
通讯作者: Sternberg, MJE
DOI: 10.1002/qsar.200310005
发表时间: 2003-07-01
期刊: QSAR & COMBINATORIAL SCIENCE
影响因子: --
作者:
Sternberg, MJE;Muggleton, SH
通讯作者: Muggleton, SH
DOI: 10.1038/73723
发表时间: 2000-03-01
影响因子: 46.9
作者:
Skolnick, J;Fetrow, JS;Kolinski, A
通讯作者: Kolinski, A
DOI: 10.1007/s10822-007-9113-3
发表时间: 2007-05-01
影响因子: 3.5
作者:
Cannon, Edward O.;Amini, Ata;Mitchell, John B. O.
通讯作者: Mitchell, John B. O.
DOI: 10.1021/ci600223d
发表时间: 2007-05-01
影响因子: 5.6
作者:
Amini, Ata;Muggleton, Stephen H.;Sternberg, Michael J. E.
通讯作者: Sternberg, Michael J. E.