Modeling urban coastal flood severity from crowd-sourced flood reports using Poisson regression and Random Forest

Modeling urban coastal flood severity from crowd-sourced flood reports using Poisson regression and Random Forest
复制标题

DOI:
10.1016/j.jhydrol.2018.01.044
复制
发表时间:
2018-04
影响因子:
6.4
通讯作者:
J. Sadler;J. Goodall;Mohamed M. Morsy;K. Spencer
J. Sadler;J. Goodall;Mohamed M. Morsy;K. Spencer
中科院分区:
地球科学1区
文献类型:
--
作者:
J. Sadler;J. Goodall;Mohamed M. Morsy;K. Spencer

文献摘要

被引文献

相似文献

海平面上升已经导致沿海洪水更加频繁和严重,而且这种趋势可能会持续下去。洪水预测是沿海城市适应和缓解这一日益严重的问题的能力的重要组成部分。然而,复杂的沿海城市水文系统并不总是适合基于物理的洪水预测方法。本文提出了一种使用数据驱动的方法来估计城市沿海环境中洪水严重程度的方法,该方法使用众包数据(一种非传统但不断增长的数据源)以及环境观测数据。泊松回归和随机森林回归这两种数据驱动模型经过训练,可以在输入大量环境数据(即降雨量、潮汐、地下水位和风况)的情况下,预测每次风暴事件的洪水报告数量,作为洪水严重程度的代理。该方法使用 2010 年 9 月至 2016 年 10 月美国弗吉尼亚州诺福克的数据进行了演示。使用质量控制的众包街道洪水报告(每次风暴事件 1 到 159 份,涉及 45 次风暴事件)来训练和评估模型。随机森林在预测洪水报告数量方面比泊松回归表现更好,并且误报率更低。根据随机森林模型,累计降雨量是迄今为止预测洪水严重程度最主要的输入变量,其次是低潮和较低的低潮。这些方法是使用数据驱动方法进行空间和时间详细的沿海城市洪水预测的第一步。
Sea level rise has already caused more frequent and severe coastal flooding and this trend will likely continue. Flood prediction is an essential part of a coastal city’s capacity to adapt to and mitigate this growing problem. Complex coastal urban hydrological systems however, do not always lend themselves easily to physically-based flood prediction approaches. This paper presents a method for using a data-driven approach to estimate flood severity in an urban coastal setting using crowd-sourced data, a non-traditional but growing data source, along with environmental observation data. Two data-driven models, Poisson regression and Random Forest regression, are trained to predict the number of flood reports per storm event as a proxy for flood severity, given extensive environmental data (i.e., rainfall, tide, groundwater table level, and wind conditions) as input. The method is demonstrated using data from Norfolk, Virginia USA from September 2010 to October 2016. Quality-controlled, crowd-sourced street flooding reports ranging from 1 to 159 per storm event for 45 storm events are used to train and evaluate the models. Random Forest performed better than Poisson regression at predicting the number of flood reports and had a lower false negative rate. From the Random Forest model, total cumulative rainfall was by far the most dominant input variable in predicting flood severity, followed by low tide and lower low tide. These methods serve as a first step toward using data-driven methods for spatially and temporally detailed coastal urban flood prediction.