Generating High Spatial Resolution Exposure Estimates from Sparse Regulatory Monitoring Data.

Generating High Spatial Resolution Exposure Estimates from Sparse Regulatory Monitoring Data.
复制标题

从稀疏的监管监测数据生成高空间分辨率暴露估计。

DOI:
10.1016/j.atmosenv.2023.120076
复制
发表时间:
2023
期刊:
Atmospheric environment (Oxford, England : 1994)
影响因子:
--
通讯作者:
Zhang,Junfeng
Zhang,Junfeng
中科院分区:
--
文献类型:
--
作者:
Ge,Yihui;Yang,Zhenchun;Lin,Yan;Hopke,PhilipK;Presto,AlbertA;Wang,Meng;Rich,DavidQ;Zhang,Junfeng

文献摘要

相似文献

随机森林算法已被广泛用于估计环境空气污染物浓度。然而,模型预测估计的准确性可能会受到与有限的测量数据相关的外推问题的影响,以训练机器学习算法。在这项研究中,我们开发和评估了两种方法,结合低成本的传感器数据,增强了随机森林模型在监测数据稀少的地区的外推能力。纽约州罗切斯特是一项怀孕队列研究的地区。获得了NAMS/SLAMS站点的PM2.5日浓度,并将其作为模型的响应变量,卫星数据、气象和土地利用变量作为预测变量。为了改进基础随机森林模型,我们使用了一个已有的低成本传感器网络的PM2.5测量值,然后进行了两步向后选择,逐步从基础模型中剔除了具有潜在排放异质性的变量。然后,我们将回归增强型随机森林方法引入到模型开发中。最后,用同期尿1-羟基芘对两种方法产生的PM2.5预测进行了评估。两步法将平均外部验证R2从0.49提高到0.65,RMSE从3.56亿μg/m~3降低到2.96亿μg/m~3。回归增强型随机森林模型的外部验证平均R2为0.54,均方根误差为3.40μg/m~3。我们还观察到两个改进模型的尿1-羟基芘水平和PM2.5预测之间的显著和可比的关系。这种PM2.5模型估计策略可以提高随机森林模型在监测数据稀疏地区的外推能力。
Random Forest algorithms have extensively been used to estimate ambient air pollutant concentrations. However, the accuracy of model-predicted estimates can suffer from extrapolation problems associated with limited measurement data to train the machine learning algorithms. In this study, we developed and evaluated two approaches, incorporating low-cost sensor data, that enhanced the extrapolating ability of random-forest models in areas with sparse monitoring data. Rochester, NY is the area of a pregnancy-cohort study. Daily PM2.5concentrations from the NAMS/SLAMS sites were obtained and used as the response variable in the model, with satellite data, meteorological, and land-use variables included as predictors. To improve the base random-forest models, we used PM2.5measurements from a pre-existing low-cost sensors network, and then conducted a two-step backward selection to gradually eliminate variables with potential emission heterogeneity from the base models. We then introduced the regression-enhanced random forest method into the model development. Finally, contemporaneous urinary 1-hydroxypyrene was used to evaluate the PM2.5predictions generated from the two approaches. The two-step approach increased the average external validation R2from 0.49 to 0.65, and decreased the RMSE from 3.56 μg/m3to 2.96 μg/m3. For the regression-enhanced random forest models, the average R2of the external validation was 0.54, and the RMSE was 3.40 μg/m3. We also observed significant and comparable relationships between urinary 1-hydroxypyrene levels and PM2.5predictions from both improved models. This PM2.5model estimation strategy could improve the extrapolating ability of random forest models in areas with sparse monitoring data.