Improved coarse-graining of Markov state models via explicit consideration of statistical uncertainty

Improved coarse-graining of Markov state models via explicit consideration of statistical uncertainty
复制标题

DOI:
10.1063/1.4755751
复制
发表时间:
2012-10-07
影响因子:
4.4
通讯作者:
Bowman, Gregory R.
Bowman, Gregory R.
中科院分区:
化学2区
文献类型:
--
作者:
Bowman, Gregory R.

文献摘要

被引文献

相似文献

马尔可夫状态模型 (MSM)(或离散时间主方程模型)是对蛋白质等分子系统的结构和功能进行建模的有效方法。不幸的是,具有足够多状态以与实验进行定量联系的 MSM(即使对于小型系统,通常也有数万个状态)通常太复杂而难以理解。在这里,我提出了一种贝叶斯凝聚聚类引擎(BACE),用于粗粒度此类马尔可夫模型,从而降低其复杂性并使它们更易于理解。该算法的一个重要特征是它能够明确解释有限采样产生的模型参数的统计不确定性。这一进展建立在最近的一些工作的基础上,这些工作强调了在 MSM 分析中考虑不确定性的重要性,并且与粗粒度马尔可夫状态模型的现有方法相比具有显着的优势。我在这里导出的用于确定要合并哪些状态的封闭式表达式相当于广义詹森-香农散度,这是信息论中与相对熵相关的一个重要度量。因此,该方法在最小化信息损失方面具有有吸引力的信息论解释。该算法自下而上的性质可能使其特别适合构建介尺度模型。我还提出了一种极其有效的贝叶斯模型比较表达式,可用于识别 BACE 模型层次结构中最有意义的级别。 (C) 2012 年美国物理研究所。 [http://dx.doi.org/10.1063/1.4755751]
Markov state models (MSMs)-or discrete-time master equation models-are a powerful way of modeling the structure and function of molecular systems like proteins. Unfortunately, MSMs with sufficiently many states to make a quantitative connection with experiments (often tens of thousands of states even for small systems) are generally too complicated to understand. Here, I present a Bayesian agglomerative clustering engine (BACE) for coarse-graining such Markov models, thereby reducing their complexity and making them more comprehensible. An important feature of this algorithm is its ability to explicitly account for statistical uncertainty in model parameters that arises from finite sampling. This advance builds on a number of recent works highlighting the importance of accounting for uncertainty in the analysis of MSMs and provides significant advantages over existing methods for coarse-graining Markov state models. The closed-form expression I derive here for determining which states to merge is equivalent to the generalized Jensen-Shannon divergence, an important measure from information theory that is related to the relative entropy. Therefore, the method has an appealing information theoretic interpretation in terms of minimizing information loss. The bottom-up nature of the algorithm likely makes it particularly well suited for constructing mesoscale models. I also present an extremely efficient expression for Bayesian model comparison that can be used to identify the most meaningful levels of the hierarchy of models from BACE. (C) 2012 American Institute of Physics. [http://dx.doi.org/10.1063/1.4755751]