FF-IR: An information retrieval system for flash flood events developed by integrating public-domain data and machine learning

FF-IR: An information retrieval system for flash flood events developed by integrating public-domain data and machine learning
复制标题

DOI:
10.1016/j.envsoft.2023.105734
复制
发表时间:
2023-09
期刊:
Environ. Model. Softw.
影响因子:
--
通讯作者:
Rohan Singh Wilkho;N. Gharaibeh;Shih-Nun Chang;Lei Zou
Rohan Singh Wilkho;N. Gharaibeh;Shih-Nun Chang;Lei Zou
中科院分区:
其他
文献类型:
--
作者:
Rohan Singh Wilkho;N. Gharaibeh;Shih-Nun Chang;Lei Zou

文献摘要

被引文献

相似文献

关于山洪(FF)事件的结构化数据库信息有限,缺乏新兴数据(例如,视觉媒体)。网络上有丰富的信息可以弥补这一差距。然而,搜索引擎返回的一长串网页列表中充满了商业和不相关的信息。为了解决这个问题,我们开发了一个FF信息检索(IR)系统(FF-IR)。该系统以新颖的方式使用机器学习(ML)模型来自动化和增强这一IR过程。FF-IR由三个步骤组成:(1)从公开可用的Storm Events数据集创建特定于事件的搜索查询,并将其引导到Google以收集候选网页;(2)将候选网页转换为相关性特征;(3)使用我们的ML模型将每个候选网页分类为相关或不相关。FF-IR比直接Google搜索的性能高出100%以上,以F2分数衡量。自然灾害研究人员和从业人员可以使用FF-IR来促进FF风险评估和减灾规划。
Structured databases on flash flood (FF) events have limited information and lack emerging data (e.g., visual media). The web is rich with information that can bridge this gap. However, search engines return long lists of webpages cluttered with commercial and irrelevant information. To address this challenge, we developed a FF information retrieval (IR) system (FF-IR). The system uses machine learning (ML) models in novel ways to automate and enhance this IR process. FF-IR consists of three steps: (1) creates event-specific search queries from the publicly available Storm Events dataset and directs them to Google to collect candidate webpages; (2) transforms the candidate webpages to relevance features; and (3) classifies each candidate webpage as relevant or non-relevant using our ML models. FF-IR outperforms direct Google searches by over 100%, measured by the F2-score. Natural hazard researchers and practitioners can use FF-IR to facilitate FF risk assessments and mitigation planning.