Accuracy of Predictions Made by Machine Learned Models for Biocrude Yields Obtained from Hydrothermal Liquefaction of Organic Wastes

Accuracy of Predictions Made by Machine Learned Models for Biocrude Yields Obtained from Hydrothermal Liquefaction of Organic Wastes
复制标题

DOI:
10.1016/j.cej.2022.136013
复制
发表时间:
2022-03
影响因子:
15.1
通讯作者:
Feng Cheng;Elizabeth R. Belden;Wenjing Li;Muntasir Shahabuddin;R. Paffenroth;M. Timko
Feng Cheng;Elizabeth R. Belden;Wenjing Li;Muntasir Shahabuddin;R. Paffenroth;M. Timko
中科院分区:
工程技术1区
文献类型:
--
作者:
Feng Cheng;Elizabeth R. Belden;Wenjing Li;Muntasir Shahabuddin;R. Paffenroth;M. Timko

文献摘要

相似文献

水热液化(HTL)具有将大量湿有机废物转化为可再生燃料的潜力。由于HTL由复杂的反应网络组成,因此其生物原油产率的确定性的、基于物理的预测极其困难。数据驱动方法提供了一种替代基于物理的方法;然而,必须进行严格的测试,以确保数据驱动方法预测的准确性。为此,收集了一个由公开文献中出现的570个数据点组成的数据集。数据集被分为训练、验证和测试子集,并用于评估不同的机器学习回归方法来预测生物原油产量。在测试的算法中,随机森林和极端梯度提升(XGBoost)预测生物原油产量的测试集,没有被用于训练具有最大的准确性,均方根误差(RMSE)分别为8.34和8.57。随机森林模型的进一步改进将其RMSE降低到8.07。相比之下,一系列文献模型的预测结果的RMSE范围从最准确的情况下的9.16到最不准确的情况下的27.6;大多数文献模型产生的RMSE值> 10。使用最准确的随机森林模型和概率经济分析的生物原油产量预测发现,模型的准确性足以根据预计的最低燃料销售价格优先分配资源。这里提出的模型和分析代表了一个重大的进步,能够使用现成的数据来预测生物原油产量的新原料,以前没有研究。
Hydrothermal liquefaction (HTL) has potential for converting abundant wet organic wastes into renewable fuels. Because HTL consists of a complex reaction network, deterministic, physics-based prediction of its biocrude yield is prohibitively difficult. Data-driven methods provide an alternative to the physics-based approach; however, rigorous testing must be performed to ensure the accuracy of predictions made by data-driven methods. To this end, a data set was assembled consisting of 570 data points appearing in the open literature. The data set was divided into training, validation, and test sub-sets and used for evaluating different machine learning regression approaches to predict biocrude yield. Among the tested algorithms, Random Forest and eXtreme Gradient Boosting (XGBoost) predicted biocrude yields in a test set that had not been used for training with the greatest accuracy, with root mean square errors (RMSE) of 8.34 and 8.57, respectively. Further refinement of the Random Forest model reduced its RMSE to 8.07. In comparison, predictions of a series of literature models resulted in RMSE ranging from 9.16 in the most accurate case to 27.6 in the least accurate; most literature models yielded RMSE values > 10. Using biocrude yield predictions from the most accurate Random Forest model and a probabilistic economic analysis found that the model accuracy is sufficient to prioritize allocation of resources based on projected minimum fuel selling price. The models and analysis presented here represent a major advance in the ability to use readily available data to predict biocrude yields on new feedstocks that have not previously been studied.