Automated Filtering Big Visual Data from Drones for Enhanced Visual Analytics in Construction

Automated Filtering Big Visual Data from Drones for Enhanced Visual Analytics in Construction
复制标题

自动过滤来自无人机的大视觉数据,以增强施工中的视觉分析

DOI:
10.1061/9780784481264.039
复制
发表时间:
2018
期刊:
ASCE Construction Research Congress 2018
影响因子:
--
通讯作者:
Ham, Youngjib
Ham, Youngjib
中科院分区:
--
文献类型:
--
作者:
Kamari, MirSalar;Ham, Youngjib

文献摘要

被引文献

相似文献

如今,为了评估和记录建筑和建筑性能,通过配备摄像头的平台(如可穿戴摄像头、无人机/地面车辆和智能手机)捕获和存储大量视觉数据。然而,由于记录这种视觉数据的不间断方式,并非所有捕获的连续片段中的帧都是有意拍摄的,因此并非每个帧都值得被处理用于施工和建筑物性能分析。由于许多帧只是具有非施工相关内容,因此在处理视觉数据之前,应根据与视觉评估目标的关联手动调查每个记录帧的内容。为了应对这些挑战,本文旨在自动过滤不需要人工注释的建筑大视觉数据。为了克服使用手动标记图像的纯判别方法的挑战,我们构建了一个生成模型与未标记的视觉数据集,并使用它来寻找施工相关的帧在大的视觉数据集从工地。首先,通过基于合成的捕捉点检测和域自适应,我们过滤和删除大部分意外记录的画面。然后,我们创建判别分类器,用来自工地的视觉数据训练,以消除非施工相关的图像。为了评估所提出的方法的可靠性,我们已经获得了地面真理的基础上,人类的判断,为我们的测试数据集中的每张照片。尽管学习没有任何显式的标签,所提出的方法显示了合理的实用范围的准确性,这通常优于以前的捕捉点检测。通过算例分析,详细讨论了该算法的保真度.通过能够专注于选择性的视觉数据,从业者将花费更少的时间浏览大量的视觉数据;而是花更多的时间来研究如何利用视觉数据来促进构建环境中的决策。
Nowadays, to assess and document construction and building performance, large amount of visual data are captured and stored through camera equipped platforms such as wearable cameras, unmanned aerial/ground vehicles, and smart phones. However, due to the nonstop fashion in recording such visual data, not all of the frames in captured consecutive footages are intentionally taken, and thus not every frame is worthy of being processed for construction and building performance analysis. Since many frames will simply have non-construction related contents, before processing the visual data, the content of each recorded frame should be manually investigated depending on the association with the goal of the visual assessment. To address such challenges, this paper aims to automatically filter construction big visual data that requires no human annotations. To overcome challenges in pure discriminative approach using manually labeled images, we construct a generative model with unlabeled visual dataset, and use it to find construction-related frames in big visual dataset from jobsites. First, through composition-based snap point detection together with domain adaptation, we filter and remove most of accidently recorded frames in the footage. Then, we create discriminative classifier trained with visual data from jobsites to eliminate non-construction related images. To evaluate the reliability of the proposed method, we have obtained the ground truth based on human judgment for each photo in our testing dataset. Despite learning without any explicit labels, the proposed method shows a reasonable practical range of accuracy, which generally outperforms prior snap point detection. Through the case studies, the fidelity of the algorithm is discussed in detail. By being able to focus on selective visual data, practitioners will spend less time on browsing large amounts of visual data; rather spend more time on looking at how to leverage the visual data to facilitate decision-makings in built environments.