A bootstrap-aggregated hybrid semi-parametric modeling framework for bioprocess development

A bootstrap-aggregated hybrid semi-parametric modeling framework for bioprocess development
复制标题

DOI:
10.1007/s00449-019-02181-y
复制
发表时间:
2019-11-01
影响因子:
3.8
通讯作者:
von Stosch, Moritz
von Stosch, Moritz
中科院分区:
工程技术3区
文献类型:
--
作者:
Pinto, Jose;de Azevedo, Cristiana Rodrigues;von Stosch, Moritz

文献摘要

被引文献

相似文献

混合半参数建模结合了机械和机器学习方法,已被证明是一种强大的流程开发方法。本文提出了引导聚合,以在通过实验统计设计获得过程数据时提高混合半参数模型的预测能力。解决了补料分批大肠杆菌优化问题,其中三个因素(生物量生长设定点、温度和诱导时的生物量浓度)经过统计设计,以确定最佳细胞生长和重组蛋白表达条件。综合数据集是使用三种不同的设计方法生成的,即 Box-Behnken、中心复合和 Doehlert 设计。为这三种设计开发了引导聚合混合模型,并与各自的非聚合版本进行了比较。结果表明,引导聚合显着降低了所有三种设计的新批次实验的预测均方误差。要聚合的(最佳)模型的数量是一个关键的校准参数,需要在每个问题中进行微调。在识别过程最佳性方面,Doehlert 设计略优于其他设计。最后,多个预测的可用性允许计算模型不同部分的误差范围,这提供了对模型组件内预测变化的额外洞察。
Hybrid semi-parametric modeling, combining mechanistic and machine-learning methods, has proven to be a powerful method for process development. This paper proposes bootstrap aggregation to increase the predictive power of hybrid semi-parametric models when the process data are obtained by statistical design of experiments. A fed-batch Escherichia coli optimization problem is addressed, in which three factors (biomass growth setpoint, temperature, and biomass concentration at induction) were designed statistically to identify optimal cell growth and recombinant protein expression conditions. Synthetic data sets were generated applying three distinct design methods, namely, Box-Behnken, central composite, and Doehlert design. Bootstrap-aggregated hybrid models were developed for the three designs and compared against the respective non-aggregated versions. It is shown that bootstrap aggregation significantly decreases the prediction mean squared error of new batch experiments for all three designs. The number of (best) models to aggregate is a key calibration parameter that needs to be fine-tuned in each problem. The Doehlert design was slightly better than the other designs in the identification of the process optimum. Finally, the availability of several predictions allowed computing error bounds for the different parts of the model, which provides an additional insight into the variation of predictions within the model components.