Bedload transport rate prediction: Application of novel hybrid data mining techniques

Bedload transport rate prediction: Application of novel hybrid data mining techniques
复制标题

DOI:
10.1016/j.jhydrol.2020.124774
复制
发表时间:
2020-06-01
影响因子:
6.4
通讯作者:
Dieu Tien Bui
Dieu Tien Bui
中科院分区:
地球科学1区
文献类型:
--
作者:
Khosravi, Khabat;Cooper, James R.;Dieu Tien Bui

文献摘要

被引文献

相似文献

准确预测砾石河床河床荷载输送仍然是河流科学中的重大挑战。然而,数据挖掘算法提供垫料运输模型的潜力仍有待探索。这项研究首次对一系列独立和混合数据挖掘模型的预测能力进行了量化。使用实验室水槽实验中收集的床荷运输数据,评估了四种最近开发的独立数据挖掘技术的性能 - M5P、随机树 (RT)、随机森林 (RF) 和减少误差修剪树 (REPT) - 以及使用装袋 (BA) 数据挖掘算法训练的四种类型的混合算法(BA-M5P、BA-RF、BA-RT 和 BA-REPT)。主要发现有四个方面。首先,BA-M5P模型的预测能力最高(R-2 = 0.943;RMSE = 0.061 kg m(-1) s(-1);MAE = 0.040 kg m(-1) s(-1);NSE = 0.945;PBIAS = -1.60),其次是M5P、BA-RT、RT、BA-RF、RF、BA-REPT和报告。除 BA-REPT 和 REPT 模型“令人满意”外,所有模型均表现出“非常好”的性能。其次,M5P、BA-RT 和 RT 模型低估了床质输送率,而 BA-M5P、BA-RF、RF、BA-REPT 和 REPT 模型高估了床质输送率。第三,流速对底土输送速率(PCC = 0.760)影响最大,其次是剪切应力(PCC = 0.709)、流量(PCC = 0.668)、底床剪切速度(PCC = 0.663)、底床坡度(PCC = 0.490)、水流深度(PCC = 0.303)、中值沉积物直径(PCC = 0.247)和相对粗糙度(PCC = 0.003)。第四,树的最大深度是基于决策树的算法中最敏感的算子,批量大小、执行槽数和小数位数对模型的预测能力没有任何影响。总体而言,结果表明,混合数据挖掘技术比独立数据挖掘模型提供了更准确的底泥输送率预测。特别是,使用 Bagging 数据挖掘算法训练的 M5P 模型具有对砾石河床河床荷载输送进行可靠预测的巨大潜力。
The accurate prediction of bedload transport in gravel-bed rivers remains a significant challenge in river science. However the potential for data mining algorithms to provide models of bedload transport have yet to be explored. This study provides the first quantification of the predictive power of a range of standalone and hybrid data mining models. Using bedload transport data collected in laboratory flume experiments, the performance of four types of recently developed standalone data mining techniques - the M5P, random tree (RT), random forest (RF) and the reduced error pruning tree (REPT) - are assessed, along with four types of hybrid algorithms trained with a Bagging (BA) data mining algorithm (BA-M5P, BA-RF, BA-RT and BA-REPT). The main findings are fourfold. First, the BA-M5P model had the highest prediction power (R-2 = 0.943; RMSE = 0.061 kg m(-1) s(-1) ; MAE = 0.040 kg m(-1) s(-1); NSE = 0.945; PBIAS = -1.60) followed by M5P, BA-RT, RT, BA-RF, RF, BA-REPT, and REPT. All models displayed 'very good' performance except the BA-REPT and REPT model, which were `satisfactory'. Second, the M5P, BA-RT, and RT models underestimated, and the BA-M5P, BA-RF, RF, BA-REPT and REPT models overestimated, bedload transport rates. Third, flow velocity had the most significant impact on bedload transport rate (PCC = 0.760) followed by shear stress (PCC = 0.709), discharge (PCC = 0.668), bed shear velocity (PCC = 0.663), bed slope (PCC = 0.490), flow depth (PCC = 0.303), median sediment diameter (PCC = 0.247), and relative roughness (PCC = 0.003). Fourth, the maximum depth of tree was the most sensitive operator in decision tree-based algorithms, and batch size, number of execution slots and number of decimal places did not have any impact on model' prediction power. Overall the results revealed that hybrid data mining techniques provide more accurate predictions of bedload transport rate than standalone data mining models. In particular, M5P models, trained with a Bagging data mining algorithm, have great potential to produce robust predictions of bedload transport in gravel-bed rivers.