Imputation of missing sub-hourly precipitation data in a large sensor network: a machine learning approach

Imputation of missing sub-hourly precipitation data in a large sensor network: a machine learning approach
复制标题

DOI:
10.1016/j.jhydrol.2020.125126
复制
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Benedict D. Chivers;J. Wallbank;S. Cole;O. Šebek;S. Stanley;M. Fry;G. Leontidis
Benedict D. Chivers;J. Wallbank;S. Cole;O. Šebek;S. Stanley;M. Fry;G. Leontidis
中科院分区:
其他
文献类型:
--
作者:
Benedict D. Chivers;J. Wallbank;S. Cole;O. Šebek;S. Stanley;M. Fry;G. Leontidis

文献摘要

相似文献

以亚小时分辨率收集的降水数据在很大程度上是随机的,并且在降雨和非降雨的持续时间上高度不平衡,因此对丢失的数据恢复提出了具体的挑战。在这里,我们提出了一个两步分析,利用当前的机器学习技术,通过将任务分解为(a)降雨或非降雨样本的分类,以及(b)回归预测降雨样本的绝对值,来推算每隔30分钟采样的降水数据。通过对英国37个气象站的调查,这种机器学习过程产生了比利用邻近雨量计的既定表面拟合技术更准确的降水数据预测。增加机器学习算法训练的可用功能,通过将目标站点的天气数据与提供最高性能的外部雨量计集成,提高了性能。这种方法通过利用同时收集的环境数据中的信息来通知机器学习模型,从而对缺失的降雨数据做出准确的预测。从弱相关变量中捕获复杂的非线性关系对于以亚小时分辨率恢复数据至关重要。这种数据恢复管道可以开发和部署,用于在高时间分辨率下对正在进行的数据集中的缺失值进行高度自动化和近乎瞬时的输入。
Precipitation data collected at sub-hourly resolution represents specific challenges for missing data recovery by being largely stochastic in nature and highly unbalanced in the duration of rain vs non-rain. Here we present a two-step analysis utilising current machine learning techniques for imputing precipitation data sampled at 30-minute intervals by devolving the task into (a) the classification of rain or non-rain samples, and (b) regressing the absolute values of predicted rain samples. Investigating 37 weather stations in the UK, this machine learning process produces more accurate predictions for recovering precipitation data than an established surface fitting technique utilising neighbouring rain gauges. Increasing available features for the training of machine learning algorithms increases performance with the integration of weather data at the target site with externally sourced rain gauges providing the highest performance. This method informs machine learning models by utilising information in concurrently collected environmental data to make accurate predictions of missing rain data. Capturing complex non-linear relationships from weakly correlated variables is critical for data recovery at sub-hourly resolutions. Such pipelines for data recovery can be developed and deployed for highly automated and near instantaneous imputation of missing values in ongoing datasets at high temporal resolutions.