Estimating Cellular Goals from High-Dimensional Biological Data

Estimating Cellular Goals from High-Dimensional Biological Data
复制标题

DOI:
10.1145/3292500.3330775
复制
发表时间:
2018-07
期刊:
Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Laurence Yang;José Bento;Jean-Christophe Lachance;B. Palsson
Laurence Yang;José Bento;Jean-Christophe Lachance;B. Palsson
中科院分区:
其他
文献类型:
--
作者:
Laurence Yang;José Bento;Jean-Christophe Lachance;B. Palsson

文献摘要

相似文献

基于优化的模型已经用于预测细胞行为超过25年。这些模型的约束来自基因组注释、测量的细胞大分子组成,以及通过测量细胞在不同条件下的生长速率和代谢。细胞目标(细胞试图解决的优化问题)对于许多生物体(包括人类或哺乳动物细胞)来说,通过实验推导具有挑战性,因为它们具有复杂的代谢能力并且尚未得到很好的理解。从数据中学习目标的现有方法包括(a)估计线性目标函数,或(b)估计模拟复杂生化反应和约束细胞操作的线性约束。后一种方法很重要,因为通常已知的反应不足以解释观察结果;因此,有必要通过学习新的反应来自动扩展模型的复杂性。然而,这会导致非凸优化问题,并且现有的工具不能扩展到实际的大型代谢模型。因此,尽管约束估计在模拟细胞代谢方面有好处,但它仍然很少被使用,这对于开发针对病原体的新型抗菌剂、发现癌症药物靶点和生产增值化学品非常重要。在这里,我们开发了第一种方法来估计约束反应的数据,可以扩展到实际的大型代谢模型。以前的工具用于解决少于75种反应和60种代谢物的问题,这限制了实际尺寸的应用。我们使用75种不同生物(包括细菌、酵母和哺乳动物)的大规模代谢网络模型进行了广泛的实验,并表明我们的算法可以恢复细胞约束反应。恢复的约束条件能够准确预测训练数据中未见的数百种生长环境中的代谢状态,即使在一些测量缺失的情况下,我们也能恢复有用的细胞目标。
Optimization-based models have been used to predict cellular behavior for over 25 years. The constraints in these models are derived from genome annotations, measured macromolecular composition of cells, and by measuring the cell's growth rate and metabolism in different conditions. The cellular goal (the optimization problem that the cell is trying to solve) can be challenging to derive experimentally for many organisms, including human or mammalian cells, which have complex metabolic capabilities and are not well understood. Existing approaches to learning goals from data include (a) estimating a linear objective function, or (b) estimating linear constraints that model complex biochemical reactions and constrain the cell's operation. The latter approach is important because often the known reactions are not enough to explain observations; therefore, there is a need to extend automatically the model complexity by learning new reactions. However, this leads to nonconvex optimization problems, and existing tools cannot scale to realistically large metabolic models. Hence, constraint estimation is still used sparingly despite its benefits for modeling cell metabolism, which is important for developing novel antimicrobials against pathogens, discovering cancer drug targets, and producing value-added chemicals. Here, we develop the first approach to estimating constraint reactions from data that can scale to realistically large metabolic models. Previous tools were used on problems having less than 75 reactions and 60 metabolites, which limits real-life-size applications. We perform extensive experiments using 75 large-scale metabolic network models for different organisms (including bacteria, yeasts, and mammals) and show that our algorithm can recover cellular constraint reactions. The recovered constraints enable accurate prediction of metabolic states in hundreds of growth environments not seen in training data, and we recover useful cellular goals even when some measurements are missing.