A mechanism-aware and multiomic machine-learning pipeline characterizes yeast cell growth

A mechanism-aware and multiomic machine-learning pipeline characterizes yeast cell growth
复制标题

DOI:
10.1073/pnas.2002959117
复制
发表时间:
2020-08-04
影响因子:
11.1
通讯作者:
Angione, Claudio
Angione, Claudio
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Culley, Christopher;Vijayakumar, Supreeta;Angione, Claudio

文献摘要

被引文献

相似文献

代谢建模和机器学习是新兴的下一代系统和合成生物学工具的关键组成部分,目标是基因型-表型-环境关系。现在越来越清楚的是,它们的价值不是单独使用,而是结合起来才能最大化。然而,整合这两种框架的潜力在很大程度上尚未得到探索。我们提出,严格评估和比较基于机器学习的数据集成技术,将基因表达谱与计算生成的代谢通量数据相结合,以预测酵母细胞生长。为此,我们为1143个酿酒酵母突变体创建了菌株特异性代谢模型,并测试了27种机器学习方法,其中包括最先进的特征选择和多视图学习方法。我们提出了一个使用通量组和转录组数据的多视图神经网络,表明前者增加了后者的预测准确性,并揭示了不能直接从基因表达中推断出来的功能模式。我们在另一个实验中生成的另外86个菌株上测试了所提出的神经网络,从而验证了其对额外独立数据集的鲁棒性。最后,我们表明,引入机制通量特征也提高了对代谢重建中未建模基因的敲除菌株的预测。因此,我们的研究结果表明,将实验线索与基于已知生物化学的计算机模型融合,可以为生物知情和可解释的机器学习提供不一致的信息。总的来说,这项研究为理解和操纵复杂表型提供了工具,提高了预测的准确性和可识别的机制生物学见解的程度。
Metabolic modeling and machine learning are key components in the emerging next generation of systems and synthetic biology tools, targeting the genotype-phenotype-environment relationship. Rather than being used in isolation, it is becoming clear that their value is maximized when they are combined. However, the potential of integrating these two frameworks for omic data augmentation and integration is largely unexplored. We propose, rigorously assess, and compare machine-learning- based data integration techniques, combining gene expression profiles with computationally generated metabolic flux data to predict yeast cell growth. To this end, we create strain-specific metabolic models for 1,143 Saccharomyces cerevisiae mutants and we test 27 machine-learning methods, incorporating stateof-the-art feature selection and multiview learning approaches. We propose a multiview neural network using fluxomic and transcriptomic data, showing that the former increases the predictive accuracy of the latter and reveals functional patterns that are not directly deducible from gene expression alone. We test the proposed neural network on a further 86 strains generated in a different experiment, therefore verifying its robustness to an additional independent dataset. Finally, we show that introducing mechanistic flux features improves the predictions also for knockout strains whose genes were not modeled in the metabolic reconstruction. Our results thus demonstrate that fusing experimental cues with in silico models, based on known biochemistry, can contribute with disjoint information toward biologically informed and interpretable machine learning. Overall, this study provides tools for understanding and manipulating complex phenotypes, increasing both the prediction accuracy and the extent of discernible mechanistic biological insights.