Ensemble approach based on bagging, boosting and stacking for short-term prediction in agribusiness time series

Ensemble approach based on bagging, boosting and stacking for short-term prediction in agribusiness time series
复制标题

DOI:
10.1016/j.asoc.2019.105837
复制
发表时间:
2020-01-01
影响因子:
8.7
通讯作者:
Coelho, Leandro dos Santos
Coelho, Leandro dos Santos
中科院分区:
计算机科学2区
文献类型:
--
作者:
Dal Molin Ribeiro, Matheus Henrique;Coelho, Leandro dos Santos

文献摘要

被引文献

相似文献

调查用于预测农产品价格的方法的准确性是一个重要的研究领域。在这方面,必须制定有效的模式。回归集成可以用于此目的。集合是一组组合模型,它们共同作用以预测具有较低误差的响应变量。面对这一点,这项工作的一般贡献是探索回归合奏的预测能力,通过比较合奏本身,以及考虑在农业综合企业领域的单一模型(参考模型)的方法来预测价格提前一个月。在这方面,使用了每月的时间序列,涉及巴西巴拉那州生产者为每袋60公斤大豆(案例研究1)和小麦(案例研究2)支付的价格。采用整体装袋(随机森林- RF)、增压(梯度增压机- GBM和极端梯度增压机- XGB)和堆垛(STACK)。采用支持向量回归机(SVR)、多层感知器神经网络(MLP)和K近邻(KNN)作为参考模型。性能指标,如平均绝对百分比误差(MAPE),均方根误差(RMSE),平均绝对误差(MAE)和均方误差(MSE)用于模型比较。Friedman和Wilcoxon符号秩检验用于评估模型的绝对百分比误差(APE)。从测试集结果的比较来看,观察到最佳集成方法的MAPE低于1%。在这种情况下,XGB/STACK(最小绝对收缩和选择算子-KNN-XGB-SVR)和RF模型分别在案例研究1和2的短期预测任务中表现出更好的性能。与参考模型相比,XGB/STACK和RF观察到更好的APE(统计学上更小)。除此之外,基于提升的方法是一致的,在两个案例研究中都提供了良好的结果。此外,根据性能的排名是:XGB,GBM,RF,STACK,MLP,SVR和KNN。可以得出结论,集成方法提供了统计上显着的收益,减少预测误差的价格序列研究。建议使用集合预测一个月前的农产品价格,因为观察到更自信的表现,这可以提高构建模型的准确性,降低决策风险。(C)2019 Elsevier B. V.版权所有。
The investigation of the accuracy of methods employed to forecast agricultural commodities prices is an important area of study. In this context, the development of effective models is necessary. Regression ensembles can be used for this purpose. An ensemble is a set of combined models which act together to forecast a response variable with lower error. Faced with this, the general contribution of this work is to explore the predictive capability of regression ensembles by comparing ensembles among themselves, as well as with approaches that consider a single model (reference models) in the agribusiness area to forecast prices one month ahead. In this aspect, monthly time series referring to the price paid to producers in the state of Parana, Brazil for a 60 kg bag of soybean (case study 1) and wheat (case study 2) are used. The ensembles bagging (random forests - RF), boosting (gradient boosting machine - GBM and extreme gradient boosting machine - XGB), and stacking (STACK) are adopted. The support vector machine for regression (SVR), multilayer perceptron neural network (MLP) and K-nearest neighbors (KNN) are adopted as reference models. Performance measures such as mean absolute percentage error (MAPE), root mean squared error (RMSE), mean absolute error (MAE), and mean squared error (MSE) are used for models comparison. Friedman and Wilcoxon signed rank tests are applied to evaluate the models' absolute percentage errors (APE). From the comparison of test set results, MAPE lower than 1% is observed for the best ensemble approaches. In this context, the XGB/STACK (Least Absolute Shrinkage and Selection Operator-KNN-XGB-SVR) and RF models showed better performance for short-term forecasting tasks for case studies 1 and 2, respectively. Better APE (statistically smaller) is observed for XGB/STACK and RF in relation to reference models. Besides that, approaches based on boosting are consistent, providing good results in both case studies. Alongside, a rank according to the performances is: XGB, GBM, RF, STACK, MLP, SVR and KNN. It can be concluded that the ensemble approach presents statistically significant gains, reducing prediction errors for the price series studied. The use of ensembles is recommended to forecast agricultural commodities prices one month ahead, since a more assertive performance is observed, which allows to increase the accuracy of the constructed model and reduce decision-making risk. (C) 2019 Elsevier B.V. All rights reserved.