Normal mode analysis of macromolecular motions in a database framework: Developing mode concentration as a useful classifying statistic

Normal mode analysis of macromolecular motions in a database framework: Developing mode concentration as a useful classifying statistic
复制标题

DOI:
10.1002/prot.10168
复制
发表时间:
2002-09-01
影响因子:
2.9
通讯作者:
Gerstein, M
Gerstein, M
中科院分区:
生物学4区
文献类型:
--
作者:
Krebs, WG;Alexandrov, V;Gerstein, M

文献摘要

被引文献

相似文献

我们在数据库框架内使用正常模式研究蛋白质运动,在大样本上确定正常模式预测观察到的运动方向的程度,并对运动分类有用。作为我们分析的起点,我们从PDB中蛋白质的一组全面的结构排列中确定了大量蛋白质灵活性的例子。每个例子都由一对蛋白质组成,考虑到它们序列的相似性,它们的结构有很大的不同。对于每一对,我们在高通量管道中进行几何比较和绝热映射插值,最终得到3814个假定运动和每个运动的标准化统计数据。然后,我们计算了这个列表中每个运动的正常模式,确定了最接近观察到的运动方向的模式的线性组合。我们整合了我们的新运动和正常模式的计算在大分子运动数据库,通过一个新的排名界面在http://molmovdb.org。基于正态模态计算和插值,我们确定了一个新的统计量,即模态集中,它与信息量的数学概念有关,描述了观察到的运动方向可以被几个模态概括的程度。利用这一统计数据,我们能够确定3814种运动中只有少数模式能够预测实际运动方向的部分。我们还研究了模式浓度与正常模式组合的相关统计数据的比较,并将其与表征蛋白质灵活性的数量(例如,最大骨干位移或移动原子的数量)相关联。最后,与运动统计相比,我们评估了模式集中自动将运动分类为各种简单类别的能力(例如,它们是否“碎片状”)。这涉及到将决策树和特征选择(特别是机器学习技术)应用于训练和测试集,这些训练和测试集源于将动作“列表”与手动分类的动作合并。蛋白质2002;48:682 - 695。(C) 2002年wiley-liss公司。
We investigated protein motions using normal modes within a database framework, determining on a large sample the degree to which normal modes anticipate the direction of the observed motion and were useful for motions classification. As a starting point for our analysis, we identified a large number of examples of protein flexibility from a comprehensive set of structural alignments of the proteins in the PDB. Each example consisted of a pair of proteins that were considerably different in structure given their sequence similarity. On each pair, we performed geometric comparisons and adiabatic-mapping interpolations in a high-throughput pipeline, arriving at a final list of 3,814 putative motions and standardized statistics for each. We then computed the normal modes of each motion in this list, determining the linear combination of modes that best approximated the direction of the observed motion. We integrated our new motions and normal mode calculations in the Macromolecular Motions Database, through a new ranking interface at http://molmovdb.org. Based on the normal mode calculations and the interpolations, we identified a new statistic, mode concentration, related to the mathematical concept of information content, which describes the degree to which the direction of the observed motion can be summarized by a few modes. Using this statistic, we were able to determine the fraction of the 3,814 motions where one could anticipate the direction of the actual motion from only a few modes. We also investigated mode concentration in comparison to related statistics on combinations of normal modes and correlated it,with quantities characterizing protein flexibility (e.g., maximum backbone displacement or number of mobile atoms). Finally, we evaluated the ability of mode concentration to automatically classify motions into a variety of simple categories (e.g., whether or not they are "fragment-like"), in comparison to motion statistics. This involved the application of decision trees and feature selection (particular machine-learning techniques) to training and testing sets derived from merging the "list" of motions with manually classified ones. Proteins 2002;48:682-695. (C) 2002Wiley-Liss,Inc.