Estimating hourly and continuous ground-level PM2.5 concentrations using an ensemble learning algorithm: The ST-stacking model

Estimating hourly and continuous ground-level PM2.5 concentrations using an ensemble learning algorithm: The ST-stacking model
复制标题

使用集成学习算法估算每小时和连续的地面 PM2.5 浓度:ST 堆积模型

DOI:
10.1016/j.atmosenv.2019.117242
复制
发表时间:
2020-02-15
影响因子:
5
通讯作者:
Du, Qingyun
Du, Qingyun
中科院分区:
环境科学与生态学2区
文献类型:
--
作者:
Feng, Luwei;Li, Yiyan;Du, Qingyun

文献摘要

被引文献

相似文献

地面细颗粒物(PM2,5)逐时和连续浓度的估算对于PM2,5污染源识别、有针对性的政策制定和人群暴露研究至关重要。然而,目前的PM2,5估计研究严重依赖于基于卫星的气溶胶光学厚度(ACID)数据,有限的过境时间的极轨道卫星,如Terra和Aqua,夜间间隙的数据从地球静止卫星,如Himawari-8,和云污染报告的两种类型的卫星挑战时空连续PM2,5浓度的估计。本文采用时空融合的方法构建PM2. 5时空特征。具体而言,多源数据,包括时空、周期、气象、植被、人为和拓扑特征,被纳入一种集成学习方法,该方法在第一级结合了极端梯度提升(XGBoost)、k-最近邻(KNN)和反向传播神经网络(BPNN)算法,并在第二级使用线性回归(LR)进行整合。考虑PM2.5时空自相关的优化叠加策略称为ST叠加模型。该模型使用2017年为中国采集的数据进行了训练、验证和测试。ST叠加模型的平均性能优于XGBoost,ICNN和BPNN模型9.27%,R2 = 0.9191。利用该模型,绘制了2017年5月11日中国大陆24小时和连续地面PM2,5浓度,并选择北京和成都的部分地区进行更详细的分析。塔克拉玛干沙漠、华北平原、四川盆地和长江平原的PM2.5浓度在这一天远高于其他地点,这与以往研究报告的长期模式基本一致。
Estimation of hourly and continuous ground-level fine particulate matter (PM2,5) concentrations is essential for PM2,5 pollution sources identifications, targeted policy development and population exposure research. However, current PM2,5 estimation studies rely heavily on satellite-based aerosol optical depth (ACID) data, and the limited transit times of polar-orbiting satellites such as Terra and Aqua, nighttime gaps in data from geostationary satellites such as Himawari-8, and cloud contamination reported for both types of satellites challenge the estimation of spatiotemporally continuous PM2,5 concentrations. In this study, spatiotemporal PM2.5 characteristic was constructed by the spatiotemporal fusion method. Specifically, multi-source data, including spatiotemporal, periodic, meteorological, vegetation, anthropogenic and topological characteristics, were incorporated into an ensemble learning method that combined extreme gradient boosting (XGBoost), k-nearest neighbour (KNN) and back-propagation neural network (BPNN) algorithms in level 1 and used linear regression (LR) for integration in level 2. The optimized stacking strategy that considered PM2.5 spatiotemporal autocorrelation was called the ST-stacking model. The model was trained, validated and tested with data acquired for China in 2017. The ST-stacking model outperformed XGBoost, ICNN and BPNN models by 9.27% on average, with an R2 = 0.9191. Using the model, the 24-h and continuous ground-level PM2,5 concentrations in mainland China on 11 May 2017 were mapped, and parts of Beijing and Chengdu were selected for more detailed analysis. The PM2.5 concentrations in Taklimakan Desert, North China Plain, Sichuan Basin and Yangtze Plain were much higher than those in other locations on this day, which was generally consistent with the long-term patterns reported in previous studies.