A Machine learning approach for Post-Disaster data curation

A Machine learning approach for Post-Disaster data curation
复制标题

用于灾后数据管理的机器学习方法

DOI:
10.1016/j.aei.2024.102427
复制
发表时间:
2024-04
期刊:
Adv. Eng. Informatics
影响因子:
--
通讯作者:
Sun Ho Ro;Yitong Li;Jie Gong
Sun Ho Ro;Yitong Li;Jie Gong
中科院分区:
其他
文献类型:
--
作者:
Sun Ho Ro;Yitong Li;Jie Gong

文献摘要

相似文献

自然灾害后采集的图像数据在结构失效取证中起着重要作用。然而,管理和管理大量的灾后图像数据是具有挑战性的。在大多数情况下,数据用户仍然必须花费大量精力从过去几十年存档的海量图像中查找和分类图像,以便研究特定类型的灾难。针对这一问题,本文提出了一种新的基于机器学习的方法来实现对大量自然灾害后图像数据的自动标注和分类。更具体地说,该方法将预先训练的计算机视觉模型和自然语言处理模型与跟踪自然灾害的本体相结合,以便于对特定类型的图像数据的搜索和查询。由此产生的过程返回具有五个主要标签和相似度分数的每个图像,根据开发的单词嵌入模型表示其内容。利用飓风哈维拍摄的地面居民楼全景图像,对所提出的方法进行了验证和精度评估。计算的主标签与人工分配的标签相比,最小平均差异为13.32%。这一通用和适应性强的解决方案为自动执行图像标记和分类任务提供了一个实用且有价值的解决方案,有可能应用于各种图像分类,并在不同的领域和行业使用。该方法的灵活性意味着它可以进行更新和改进,以满足各个领域不断发展的需求,使其成为未来研究和开发的宝贵资产。
Image data collected after natural disasters play an important role in the forensics of structure failures. However, curating and managing large amounts of post-disaster imagery data is challenging. In most cases, data users still have to spend much effort to find and sort images from the massive amounts of images archived for past decades in order to study specific types of disasters. This paper proposes a new machine learning based approach for automating the labeling and classification of large volumes of post-natural disaster image data to address this issue. More specifically, the proposed method couples pre-trained computer vision models and a natural language processing model with an ontology tailed to natural disasters to facilitate the search and query of specific types of image data. The resulting process returns each image with five primary labels and similarity scores, representing its content based on the developed word-embedding model. Validation and accuracy assessment of the proposed methodology was conducted with ground-level residential building panoramic images from Hurricane Harvey. The computed primary labels showed a minimum average difference of 13.32% when compared to manually assigned labels. This versatile and adaptable solution offers a practical and valuable solution for automating image labeling and classification tasks, with the potential to be applied to various image classifications and used in different fields and industries. The flexibility of the method means that it can be updated and improved to meet the evolving needs of various domains, making it a valuable asset for future research and development.