Overcoming the Heuristic Nature of k-Means Clustering: Identification and Characterization of Binding Modes from Simulations of Molecular Recognition Complexes

Overcoming the Heuristic Nature of k-Means Clustering: Identification and Characterization of Binding Modes from Simulations of Molecular Recognition Complexes
复制标题

DOI:
10.1021/acs.jcim.9b01137
复制
发表时间:
2020-06-22
影响因子:
5.6
通讯作者:
Sorin, Eric J.
Sorin, Eric J.
中科院分区:
化学2区
文献类型:
--
作者:
Bremer, Parker Ladd;De Boer, Danna;Sorin, Eric J.

文献摘要

被引文献

相似文献

准确和可重复的检测和描述的热力学状态的计算数据是一个重要的问题,特别是当状态的数量是未知的先验和大型,灵活的化学系统和复杂。为此,我们报告了一种新的聚类协议,它结合了高分辨率的结构表示,蛮力重复聚类,聚类统计和优化,可重复地确定存在于数据集(k)的模拟合奏的丁酰胆碱酯酶与两个以前研究的有机磷酸盐抑制剂的除尘器的数量。我们模拟的集合中的每个结构都被描述为一个高维向量,其中的分量由化学基团水平上的特定蛋白质-抑制剂接触定义,而这些分量的大小由其各自的成对原子接触程度定义,从而允许算法区分不同程度的相互作用。这些表面加权的相互作用指纹表中的每一个超过100万的结构,从超过100亩的全原子分子动力学模拟每个复杂的,并作为输入重复的k-均值聚类。最小化的集群人口方差和范围提供了准确和可重复的识别k,从而允许从分子模拟数据的接触表的形式,简洁地封装所观察到的分子间接触基序的离散结合模式的表征。虽然本文提出的协议,以确定k和实现非启发式聚类的大规模原子模拟的数据证明,我们的方法是可推广到其他数据类型和聚类算法,是易于处理的有限的计算资源。
The accurate and reproducible detection and description of thermodynamic states in computational data is a nontrivial problem, particularly when the number of states is unknown a priori and for large, flexible chemical systems and complexes. To this end, we report a novel clustering protocol that combines high-resolution structural representation, brute-force repeat clustering, and optimization of clustering statistics to reproducibly identify the number of dusters present in a data set (k) for simulated ensembles of butyrylcholinesterase in complex with two previously studied organophosphate inhibitors. Each structure within our simulated ensembles was depicted as a high-dimensionality vector with components defined by specific protein-inhibitor contacts at the chemical group level and the magnitudes of these components defined by their respective extents of pair-wise atomic contact, thus allowing for algorithmic differentiation between varying degrees of interaction. These surface-weighted interaction fingerprints were tabulated for each of over 1 million structures from more than 100 mu s of all-atom molecular dynamics simulation per complex and used as the input for repetitive k-means clustering. Minimization of cluster population variance and range afforded accurate and reproducible identification of k, thereby allowing for the characterization of discrete binding modes from molecular simulation data in the form of contact tables that concisely encapsulate the observed intermolecular contact motifs. While the protocol presented herein to determine k and achieve non-heuristic clustering is demonstrated on data from massive atomistic simulation, our approach is generalizable to other data types and clustering algorithms, and is tractable with limited computational resources.