Application of random forest regression to the calculation of gas-phase chemistry within the GEOS-Chem chemistry model v10

Application of random forest regression to the calculation of gas-phase chemistry within the GEOS-Chem chemistry model v10
复制标题

DOI:
10.5194/gmd-12-1209-2019
复制
发表时间:
2018-10
影响因子:
5.1
通讯作者:
C. Keller;M. Evans
C. Keller;M. Evans
中科院分区:
地球科学2区
文献类型:
--
作者:
C. Keller;M. Evans

文献摘要

被引文献

相似文献

抽象的。大气化学模型是研究化学成分对环境、植被和人类健康影响的核心工具。这些模型在数值上是密集的,以前试图降低化学求解器的数值成本并没有带来变革性的变化。我们在这里展示了机器学习(在这种情况下是随机森林回归)替代大气化学传输模型中气相化学的潜力。我们的训练数据包括1个月(2013年7月)的化学条件输出以及由GEOS-Chem化学模型v10产生的模型物理状态。从这个数据集,我们训练随机森林回归模型来预测积分器后,每个传输的物种的浓度,基于物理和化学条件积分器之前。预测类型的选择对回归模型的技巧有很大的影响。我们发现最好的结果,从预测的浓度变化为长寿的物种和短寿命的物种的绝对浓度。我们还发现,从一个简单的实施化学家庭(NOx = NO + NO2)的改进。然后,我们将训练好的随机森林预测器重新应用到GEOS-Chem中,以取代数值积分器。机器学习驱动的GEOS-Chem模型与标准模拟相比表现良好。对于臭氧(O3),使用随机森林的误差(与参考模拟相比)增长缓慢,5天后的归一化平均偏差(NMB),均方根误差(RMSE)和R2分别为4.2%,35%和0.9,30天后的误差分别增加到13%,67%和0.75。在偏远地区,如热带太平洋,偏差变得最大,在那里,化学误差可以积累,几乎没有排放或沉积的平衡影响。在过污染区,模型误差小于10%,并在以下的时间序列的全模型具有显着的保真度。模拟的氮氧化物显示出类似的特征,最显著的误差发生在远离最近排放的偏远地区。对于其他物种如无机溴物种和短寿命氮物种,误差变大,NMB、RMSE和R2分别达到> 2100%、> 400%和<0.1。这种概念验证的实现比微分方程的直接积分多花费1.8倍的时间,但是优化和软件工程应该允许速度的大幅提高。我们讨论了潜在的改进,在实施中,它的一些优点,从软件和硬件的角度来看,它的局限性,以及它的适用性业务的空气质量活动。
Abstract. Atmospheric chemistry models are a central tool to study the impact of chemical constituents on the environment, vegetation and human health. These models are numerically intense, and previous attempts to reduce the numerical cost of chemistry solvers have not delivered transformative change. We show here the potential of a machine learning (in this case random forest regression) replacement for the gas-phase chemistry in atmospheric chemistry transport models. Our training data consist of 1 month (July 2013) of output of chemical conditions together with the model physical state, produced from the GEOS-Chem chemistry model v10. From this data set we train random forest regression models to predict the concentration of each transported species after the integrator, based on the physical and chemical conditions before the integrator. The choice of prediction type has a strong impact on the skill of the regression model. We find best results from predicting the change in concentration for long-lived species and the absolute concentration for short-lived species. We also find improvements from a simple implementation of chemical families (NOx = NO + NO2). We then implement the trained random forest predictors back into GEOS-Chem to replace the numerical integrator. The machine-learning-driven GEOS-Chem model compares well to the standard simulation. For ozone (O3), errors from using the random forests (compared to the reference simulation) grow slowly and after 5 days the normalized mean bias (NMB), root mean square error (RMSE) and R2 are 4.2 %, 35 % and 0.9, respectively; after 30 days the errors increase to 13 %, 67 % and 0.75, respectively. The biases become largest in remote areas such as the tropical Pacific where errors in the chemistry can accumulate with little balancing influence from emissions or deposition. Over polluted regions the model error is less than 10 % and has significant fidelity in following the time series of the full model. Modelled NOx shows similar features, with the most significant errors occurring in remote locations far from recent emissions. For other species such as inorganic bromine species and short-lived nitrogen species, errors become large, with NMB, RMSE and R2 reaching >2100 % >400 % and <0.1, respectively. This proof-of-concept implementation takes 1.8 times more time than the direct integration of the differential equations, but optimization and software engineering should allow substantial increases in speed. We discuss potential improvements in the implementation, some of its advantages from both a software and hardware perspective, its limitations, and its applicability to operational air quality activities.