Automated Machine Learning to Evaluate the Information Content of Tropospheric Trace Gas Columns for Fine Particle Estimates Over India: A Modeling Testbed

Automated Machine Learning to Evaluate the Information Content of Tropospheric Trace Gas Columns for Fine Particle Estimates Over India: A Modeling Testbed
复制标题

DOI:
10.1029/2022ms003099
复制
发表时间:
2023-03
影响因子:
6.8
通讯作者:
Zhonghua Zheng;A. Fiore;D. Westervelt;G. Milly;J. Goldsmith;A. Karambelas;G. Curci;C. Randles;Antonio R. Paiva;Chi Wang;Qingyun Wu;S. Dey
Zhonghua Zheng;A. Fiore;D. Westervelt;G. Milly;J. Goldsmith;A. Karambelas;G. Curci;C. Randles;Antonio R. Paiva;Chi Wang;Qingyun Wu;S. Dey
中科院分区:
地球科学2区
文献类型:
--
作者:
Zhonghua Zheng;A. Fiore;D. Westervelt;G. Milly;J. Goldsmith;A. Karambelas;G. Curci;C. Randles;Antonio R. Paiva;Chi Wang;Qingyun Wu;S. Dey

文献摘要

相似文献

印度很大程度上缺乏高质量且可靠的细颗粒物 (PM2.5) 实地测量。地面 PM2.5 浓度是根据公开的卫星气溶胶光学深度 (AOD) 产品结合其他信息估算得出的。先前的研究在很大程度上忽视了利用卫星对对流层微量气柱进行反演来获得额外的准确性和了解PM来源的可能性。我们使用自动机器学习 (AutoML) 方法在建模测试台中评估印度对流层微量气体柱的信息内容,以估计印度的 PM2.5,该方法根据数据集从不同机器学习工具的菜单中进行选择。然后,我们量化对流层痕量气柱、AOD、气象场和排放量的相对信息内容,以在日和月时间尺度上估算印度四个次区域的 PM2.5。我们的研究结果表明,无论具体的机器学习模型假设如何,纳入痕量气体建模柱都可以改善 PM2.5 估计值。我们使用 AutoML 算法生成的排名分数和 Spearman 排名相关性来推断或关联 PM2.5 的主要来源与次要来源可能的相对重要性,作为估计颗粒成分的第一步。我们将 AutoML 衍生模型与选定的基准机器学习模型进行比较,结果表明 AutoML 至少与用户选择的模型一样好。这项工作中使用的理想化伪观测(化学传输模型模拟)为应用对流层痕量气体的卫星反演来估计印度的细颗粒浓度奠定了基础,并阐明了 AutoML 在大气和环境研究中应用的前景。
India is largely devoid of high‐quality and reliable on‐the‐ground measurements of fine particulate matter (PM2.5). Ground‐level PM2.5 concentrations are estimated from publicly available satellite Aerosol Optical Depth (AOD) products combined with other information. Prior research has largely overlooked the possibility of gaining additional accuracy and insights into the sources of PM using satellite retrievals of tropospheric trace gas columns. We evaluate the information content of tropospheric trace gas columns for PM2.5 estimates over India within a modeling testbed using an Automated Machine Learning (AutoML) approach, which selects from a menu of different machine learning tools based on the data set. We then quantify the relative information content of tropospheric trace gas columns, AOD, meteorological fields, and emissions for estimating PM2.5 over four Indian sub‐regions on daily and monthly time scales. Our findings suggest that, regardless of the specific machine learning model assumptions, incorporating trace gas modeled columns improves PM2.5 estimates. We use the ranking scores produced from the AutoML algorithm and Spearman’s rank correlation to infer or link the possible relative importance of primary versus secondary sources of PM2.5 as a first step toward estimating particle composition. Our comparison of AutoML‐derived models to selected baseline machine learning models demonstrates that AutoML is at least as good as user‐chosen models. The idealized pseudo‐observations (chemical‐transport model simulations) used in this work lay the groundwork for applying satellite retrievals of tropospheric trace gases to estimate fine particle concentrations in India and serve to illustrate the promise of AutoML applications in atmospheric and environmental research.