Multitask Learning Method for Detecting the Visual Focus of Attention of Construction Workers

Multitask Learning Method for Detecting the Visual Focus of Attention of Construction Workers
复制标题

DOI:
10.1061/(asce)co.1943-7862.0002071
复制
发表时间:
2021-07-01
影响因子:
5.1
通讯作者:
Cai, Hubo
Cai, Hubo
中科院分区:
工程技术2区
文献类型:
--
作者:
Cai, Jiannan;Yang, Liu;Cai, Hubo

文献摘要

被引文献

相似文献

建筑工人的视觉注意力焦点(VFOA)是识别实体交互的关键线索,这反过来又有助于解释工人的意图、预测动作和理解工地环境。建筑监控摄像头的使用越来越多,提供了一种从信息丰富的图像中估算工人 VFOA 的经济高效的方法。然而,这些图像的低分辨率对检测面部特征和注视方向提出了巨大的挑战。认识到身体和头部方向为推断工人的 VFOA 提供了强有力的提示,本研究建议将 VFOA 表示为身体方向、身体姿势、头部偏航和头部俯仰的集合,并设计一个基于卷积神经网络 (CNN) 的多任务学习 (MTL) 框架,以使用低分辨率施工图像自动估计工人的 VFOA。该框架由两个模块组成。在第一个模块中,使用更快的区域 CNN (R-CNN) 对象检测器来检测和提取工人的全身图像,生成的全身图像作为第二个模块中 CNN-MTL 模型的单个输入。在第二个模块中,VFOA 估计被表述为多任务图像分类问题,其中四个分类任务(身体方向、身体姿势、头部偏航和头部俯仰)由新设计的 CNN-MTL 模型联合学习。施工视频用于训练和测试所提出的框架。结果表明,所提出的 CNN-MTL 模型在身体方向、身体姿势、头部偏航和头部俯仰分类方面分别达到了 0.91、0.95、0.86 和 0.83 的准确度。与传统的单任务学习相比,MTL方法在不影响准确性的情况下减少了近50%的训练时间。 (C) 2021 年美国土木工程师学会。
The visual focus of attention (VFOA) of construction workers is a critical cue for recognizing entity interactions, which in turn facilitates the interpretation of workers' intentions, the prediction of movements, and the comprehension of the jobsite context. The increasing use of construction surveillance cameras provides a cost-efficient way to estimate workers' VFOA from information-rich images. However, the low resolution of these images poses a great challenge to detecting the facial features and gaze directions. Recognizing that body and head orientations provide strong hints to infer workers' VFOA, this study proposes to represent the VFOA as a collection of body orientations, body poses, head yaws, and head pitches and designs a convolutional neural network (CNN)-based multitask learning (MTL) framework to automatically estimate workers' VFOA using low-resolution construction images. The framework is composed of two modules. In the first module, a Faster regional CNN (R-CNN) object detector is used to detect and extract workers' full-body images, and the resulting full-body images serve as a single input to the CNN-MTL model in the second module. In the second module, the VFOA estimation is formulated as a multitask image classification problem where four classification tasks-body orientation, body pose, head yaw, and head pitch-are jointly learned by the newly designed CNN-MTL model. Construction videos were used to train and test the proposed framework. The results show that the proposed CNN-MTL model achieves an accuracy of 0.91, 0.95, 0.86, and 0.83 in body orientation, body pose, head yaw, and head pitch classification, respectively. Compared with the conventional single-task learning, the MTL method reduces training time by almost 50% without compromising accuracy. (C) 2021 American Society of Civil Engineers.