Improving the prediction of an atmospheric chemistry transport model using gradient-boosted regression trees

Improving the prediction of an atmospheric chemistry transport model using gradient-boosted regression trees
复制标题

DOI:
10.5194/acp-2019-753
复制
发表时间:
2019-10
影响因子:
6.3
通讯作者:
P. Ivatt;M. J. Evans
P. Ivatt;M. J. Evans
中科院分区:
地球科学1区
文献类型:
--
作者:
P. Ivatt;M. J. Evans

文献摘要

被引文献

相似文献

抽象。基于过程的环境系统模型的预测是有偏见的,由于其投入和参数化的不确定性,降低了其效用。我们开发了一个预测的偏差对流层臭氧(O3,一个关键的污染物)计算的大气化学传输模型(GEOS-Chem)的基础上,从表面(EPA,EMEP和GAW)和臭氧探测网络的臭氧模型和观测的输出。我们训练了一个梯度提升决策树算法(XGBoost)来预测模型偏差(模型除以观测值),使用2010-2015年的模型和观测数据,然后我们使用2016-2017年来测试该方法。我们表明,偏差校正模型的性能大大优于未校正的模型。均方根误差从16.2 ppb降至7.5 ppb,归一化平均偏倚从0.28降至-0.04,Pearson's R从0.48增至0.84。与来自NASA ATom飞行的观测结果(不包括在训练中)的比较也显示出改善,但程度较小,将均方根误差(RMSE)从12.1减少到10.5 ppb,将归一化平均偏差(NMB)从0.08减少到0.06,并将Pearson R从0.76增加到0.79。我们将较小的改进归因于对大部分偏远对流层缺乏常规观测限制。我们证明了该方法对训练数据量的变化具有鲁棒性,需要大约一年的数据才能产生有用的性能。数据否认实验(从算法训练中删除观测站点)表明,来自一个位置(例如欧洲)的信息可以减少其他位置(例如北美)的模型偏差,这可能会提供对控制模型偏差的过程的见解。我们探讨了预测因子的选择(偏倚预测与直接预测),并得出结论,两者都可能具有实用性。我们的结论是,将机器学习方法与基于过程的模型相结合,可以为改进这些模型提供有用的工具。
Abstract. Predictions from process-based models of environmental systems are biased, due to uncertainties in their inputs and parameterizations, reducing their utility. We develop a predictor for the bias in tropospheric ozone (O3, a key pollutant) calculated by an atmospheric chemistry transport model (GEOS-Chem), based on outputs from the model and observations of ozone from both the surface (EPA, EMEP, and GAW) and the ozone-sonde networks. We train a gradient-boosted decision tree algorithm (XGBoost) to predict model bias (model divided by observation), with model and observational data for 2010–2015, and then we test the approach using the years 2016–2017. We show that the bias-corrected model performs considerably better than the uncorrected model. The root-mean-square error is reduced from 16.2 to 7.5 ppb, the normalized mean bias is reduced from 0.28 to −0.04, and Pearson's R is increased from 0.48 to 0.84. Comparisons with observations from the NASA ATom flights (which were not included in the training) also show improvements but to a smaller extent, reducing the root-mean-square error (RMSE) from 12.1 to 10.5 ppb, reducing the normalized mean bias (NMB) from 0.08 to 0.06, and increasing Pearson's R from 0.76 to 0.79. We attribute the smaller improvements to the lack of routine observational constraints for much of the remote troposphere. We show that the method is robust to variations in the volume of training data, with approximately a year of data needed to produce useful performance. Data denial experiments (removing observational sites from the algorithm training) show that information from one location (for example Europe) can reduce the model bias over other locations (for example North America) which might provide insights into the processes controlling the model bias. We explore the choice of predictor (bias prediction versus direct prediction) and conclude both may have utility. We conclude that combining machine learning approaches with process-based models may provide a useful tool for improving these models.