Pocket Similarity: Are α Carbons Enough?

Pocket Similarity: Are α Carbons Enough?
复制标题

DOI:
10.1021/ci100210c
复制
发表时间:
2010-08-01
影响因子:
5.6
通讯作者:
Labute, Paul
Labute, Paul
中科院分区:
化学2区
文献类型:
--
作者:
Feldman, Howard J.;Labute, Paul

文献摘要

被引文献

相似文献

设计了一种测量蛋白质口袋相似性的新方法,仅使用口袋残基的a碳位置。使用穷举三维C α共同子集搜索成对比较口袋,根据理化性质对残基进行分组。每次命中需要至少五个C α匹配,并且对应点之间的距离拟合到极值分布,从而产生任何给定叠加的概率分数或可能性。一组来自13个不同蛋白质家族的85个结构仅基于结合位点使用该评分进行聚类。它也被成功地用于聚类25激酶到一些亚家族。使用测试激酶查询来检索其他激酶口袋,发现使用适当的截止分数可以实现99.2%的特异性和97.5%的灵敏度。在单个3.4 GHz CPU上搜索整个蛋白质数据库(133 800个口袋)需要2到10分钟,这取决于返回的命中数。
A novel method for measuring protein pocket similarity was devised, using only the a carbon positions of the pocket residues. Pockets were compared pairwise using an exhaustive three-dimensional C alpha common subset search, grouping residues by physicochemical properties. At least five C alpha matches were required for each hit, and distances between corresponding points were fit to an Extreme Value Distribution resulting in a probabilistic score or likelihood for any given superposition. A set of 85 structures from 13 diverse protein families was clustered based on binding sites alone, using this score. It was also successfully used to cluster 25 kinases into a number of subfamilies. Using a test kinase query to retrieve other kinase pockets, it was found that a specificity of 99.2% and sensitivity of 97.5% could be achieved using an appropriate cutoff score. The search itself took from 2 to 10 min on a single 3.4 GHz CPU to search the entire Protein Data Bank (133 800 pockets), depending on the number of hits returned.