BOFdat: Generating biomass objective functions for genome-scale metabolic models from experimental data

BOFdat: Generating biomass objective functions for genome-scale metabolic models from experimental data
复制标题

DOI:
10.1371/journal.pcbi.1006971
复制
发表时间:
2019-04-01
影响因子:
4.3
通讯作者:
Jacques, Pierre-Etienne
Jacques, Pierre-Etienne
中科院分区:
生物学2区
文献类型:
--
作者:
Lachance, Jean-Christophe;Lloyd, Colton J.;Jacques, Pierre-Etienne

文献摘要

被引文献

相似文献

基因组规模代谢模型 (GEM) 是数学结构的代谢知识库,可根据基因组信息提供表型预测。 GEM 引导的生长表型预测依赖于生物量目标函数 (BOF) 的准确定义,该函数旨在包括关键的细胞生物量成分,例如主要大分子(DNA、RNA、蛋白质)、脂质、辅酶、无机离子和物种特异性成分。尽管它很重要,但目前还没有标准化的计算平台可以以数据驱动、公正的方式生成物种特定的生物量目标函数。为了填补代谢建模软件生态系统中的这一空白,我们实施了 BOFdat,这是一个用于根据实验数据定义生物质目标函数的 Python 包。 BOFdat 采用模块化实现,将 BOF 定义过程分为三个独立模块,此处定义为步骤:1) 计算主要大分子的系数,2) 识别辅酶和无机离子并估计其化学计量系数,3) 从实验数据中以无偏差的方式通过算法提取剩余的物种特异性代谢生物质前体。我们使用 BOFdat 重建了大肠杆菌模型 iML1515 的 BOF,这是该领域的黄金标准。与其他方法相比,BOFdat 生成的 BOF 具有最一致的生物量组成、生长速率和基因必要性预测准确性。 BOFdat 的安装说明可在文档中找到,源代码可在 GitHub (https://github.com/jclachance/BOFdat) 上找到。 作者摘要 通过基因组规模模型 (GEM) 制定表型预测取决于指定的目标。生物量目标函数 (BOF) 的想法是代表细胞倍增所需的所有代谢物,因此优化 BOF 相当于优化生长。了解定性和定量生物体的组成(即哪些代谢物是生长所必需的以及其比例)对于准确预测至关重要。我们实施 BOFdat 的想法是实验数据应该驱动生物质成分的定义。随着组学数据集变得越来越可用,整合它们以获得特定条件的生物量组成的可能性已成为可能,因此也是 BOFdat 的主要特征之一。虽然主要大分子、辅酶和无机离子是跨物种普遍存在的成分,但细胞中存在一些物种特异性成分,应在 BOF 中予以考虑。为了识别这些问题,我们实施了一种方法,可以最大限度地减少实验重要性数据和 GEM 驱动的预测之间的误差。因此,BOFdat 提供了一种公正的、数据驱动的方法来定义 BOF,有可能提高新基因组规模模型的质量,并大大减少生成新重建所需的时间。
Genome-scale metabolic models (GEMs) are mathematically structured knowledge bases of metabolism that provide phenotypic predictions from genomic information. GEM-guided predictions of growth phenotypes rely on the accurate definition of a biomass objective function (BOF) that is designed to include key cellular biomass components such as the major macromolecules (DNA, RNA, proteins), lipids, coenzymes, inorganic ions and species-specific components. Despite its importance, no standardized computational platform is currently available to generate species-specific biomass objective functions in a data-driven, unbiased fashion. To fill this gap in the metabolic modeling software ecosystem, we implemented BOFdat, a Python package for the definition of a Biomass Objective Function from experimental data. BOFdat has a modular implementation that divides the BOF definition process into three independent modules defined here as steps: 1) the coefficients for major macromolecules are calculated, 2) coenzymes and inorganic ions are identified and their stoichiometric coefficients estimated, 3) the remaining species-specific metabolic biomass precursors are algorithmically extracted in an unbiased way from experimental data. We used BOFdat to reconstruct the BOF of the Escherichia coli model iML1515, a gold standard in the field. The BOF generated by BOFdat resulted in the most concordant biomass composition, growth rate, and gene essentiality prediction accuracy when compared to other methods. Installation instructions for BOFdat are available in the documentation and the source code is available on GitHub (https://github.com/jclachance/BOFdat).Author summary The formulation of phenotypic predictions by genome-scale models (GEMs) is dependent on the specified objective. The idea of a biomass objective function (BOF) is to represent all metabolites necessary for cells to double so that optimizing the BOF is equivalent to optimizing growth. Knowledge of the qualitative and quantitative organism's composition (i.e. which metabolites are necessary for growth and in what proportion) is critical for accurate predictions. We implemented BOFdat with the idea that experimental data should drive the definition of the biomass composition. As omic datasets become more available, the possibility of integrating them to obtain a condition-specific biomass composition is in reach and therefore one of the main features of BOFdat. While major macromolecules, coenzymes, and inorganic ions are ubiquitous components across species, several species-specific components exist in the cell that should be accounted for in the BOF. To identify these, we implemented an approach that minimizes the error between experimental essentiality data and GEM-driven prediction. Hence BOFdat provides an unbiased, data-driven approach to defining BOF that has the potential to improve the quality of new genome-scale models and greatly decrease the time required to generate a new reconstruction.