Statistical inference in ensemble modeling of cellular metabolism

Statistical inference in ensemble modeling of cellular metabolism
复制标题

DOI:
10.1371/journal.pcbi.1007536
复制
发表时间:
2019-12-01
影响因子:
4.3
通讯作者:
Hatzimanikatis, Vassily
Hatzimanikatis, Vassily
中科院分区:
生物学2区
文献类型:
--
作者:
Hameri, Tuure;Boldi, Marc-Olivier;Hatzimanikatis, Vassily

文献摘要

被引文献

相似文献

代谢的动力学模型可以被构建来预测细胞调节和设计代谢工程策略,并且近年来已经为此开发了各种有前途的计算工作流。由于建立动力学模型所需的动力学参数值的不确定性,这些工作流程依赖于用于采样和建立描述观察到的生理学的模型群体的系综建模(EM)原理。从代谢控制分析(MCA)的动力学模型的灵敏度系数可以提供重要的洞察细胞控制周围的一个给定的生理稳态。然而,尽管考虑群体的动力学模型和它们的模型输出,目前的方法不提供足够的工具进行统计推断。为了从模型输出(如MCA敏感性系数)中得出结论,有必要对变量群体进行排序/比较。目前现有的工作流程考虑的置信区间(CI)是为每个可比变量独立推导的。因此,重要的是要为我们希望排名/比较的变量导出同时CI。在此,我们使用了大肠杆菌代谢的现有大规模动力学模型来展示单变量CI如何导致不正确的结论,并且我们提出了一个新的工作流程,该工作流程应用了三种不同的多变量统计方法。我们使用Bonferroni和精确的正常方法来建立对称的CI使用正常的假设。然后,我们建议如何自举可以计算不对称CI,同时放松这个正常的假设。我们得出结论,Bonferroni和精确正态方法可以提供简单而有效的方法来构建可靠的CI,与精确正态方法相比,Bonferroni更受青睐时,比较变量存在依赖关系。Bootstrapping,尽管其显着更高的计算成本,建议在比较变量的非正态分布。此外,我们展示了如何Bonferroni方法可以很容易地被用来估计所需的样本数,以达到一定的CI大小。作者摘要由于各种来源的不确定性,人口的代谢动力学模型通常被构造来得出结论的动态模拟生理。然而,用于构建动力学模型群体的计算集成建模(EM)框架并没有系统地处理基于模型的结论的不确定性。尽管动力学模型可以表明,改变某些酶的水平平均会增加感兴趣的通量,但如果我们不知道模型预测的确定性,那么这些信息是不完整的。相反,我们可以使用统计推断方法来量化某些模型结论位于置信区间内的置信水平。我们演示了如何三种统计方法可以应用于构建来自动力学模型的人口的输出变量分布的置信区间。我们讨论了应用这三种方法的优点和缺点,并提供建议,他们的使用EM。这将通过帮助工程师专注于最可靠的结论来改善代谢工程决策。
Kinetic models of metabolism can be constructed to predict cellular regulation and devise metabolic engineering strategies, and various promising computational workflows have been developed in recent years for this. Due to the uncertainty in the kinetic parameter values required to build kinetic models, these workflows rely on ensemble modeling (EM) principles for sampling and building populations of models describing observed physiologies. Sensitivity coefficients from metabolic control analysis (MCA) of kinetic models can provide important insight about cellular control around a given physiological steady state. However, despite considering populations of kinetic models and their model outputs, current approaches do not provide adequate tools for statistical inference. To derive conclusions from model outputs, such as MCA sensitivity coefficients, it is necessary to rank/compare populations of variables with each other. Currently existing workflows consider confidence intervals (CIs) that are derived independently for each comparable variable. Hence, it is important to derive simultaneous CIs for the variables that we wish to rank/compare. Herein, we used an existing large-scale kinetic model of Escherichia Coli metabolism to present how univariate CIs can lead to incorrect conclusions, and we present a new workflow that applies three different multivariate statistical approaches. We use the Bonferroni and the exact normal methods to build symmetric CIs using the normality assumptions. We then suggest how bootstrapping can compute asymmetric CIs whilst relaxing this normality assumption. We conclude that the Bonferroni and the exact normal methods can provide simple and efficient ways for constructing reliable CIs, with the exact normal method favored over the Bonferroni when the compared variables present dependencies. Bootstrapping, despite its significantly higher computational cost, is recommended when comparing non-normal distributions of variables. Additionally, we show how the Bonferroni method can readily be used to estimate required sample numbers to attain a certain CI size.Author summary Due to various sources of uncertainty, populations of kinetic models of metabolism are generally constructed to derive conclusions about the dynamics of the modeled physiology. However, computational ensemble modeling (EM) frameworks for building populations of kinetic models do not systematically handle the uncertainty underlying model-based conclusions. Although kinetic models could suggest that modifying the level of certain enzymes would on average increase a flux of interest, this information is incomplete if we do not know with what certainty the model predicts this. We can instead use statistical inference approaches to quantify the level of confidence for which certain model conclusions lie within the confidence intervals. We demonstrate how three statistical methodologies can be applied to construct confidence intervals around distributions of output variables derived from populations of kinetic models. We discuss the advantages and disadvantages of applying these three methods and provide advice on their usage in EM. This will lead to improved metabolic engineering decisions by helping engineers focus on the most robust conclusions.