Multi-Modal Recognition of Worker Activity for Human-Centered Intelligent Manufacturing

Multi-Modal Recognition of Worker Activity for Human-Centered Intelligent Manufacturing
复制标题

DOI:
10.1016/j.engappai.2020.103868
复制
发表时间:
2019-08
期刊:
Eng. Appl. Artif. Intell.
影响因子:
--
通讯作者:
Wenjin Tao;M. Leu;Zhaozheng Yin
Wenjin Tao;M. Leu;Zhaozheng Yin
中科院分区:
其他
文献类型:
--
作者:
Wenjin Tao;M. Leu;Zhaozheng Yin

文献摘要

被引文献

相似文献

本研究旨在感知和理解以人为中心的智能制造系统中工人的活动。我们提出了一种新的多模态方法,通过利用来自不同传感器和不同模式的信息来识别工人的活动。具体而言,采用智能臂带和视觉摄像头分别捕获惯性测量单元(IMU)信号和视频。对于IMU信号,我们在频率域和空间域设计了两种新的特征转换机制,将捕获的IMU信号组装为图像,从而允许使用卷积神经网络学习最具判别性的特征。除了上述两种模式,我们还提出了视频数据的另外两种模式,即在视频帧和视频剪辑级别。四种模式中的每一种都返回活动预测的概率分布。然后,将这些概率分布进行融合,输出工人活动分类结果。建立了一个工人活动数据集,该数据集目前包含6种装配任务中常见的活动,即抓取工具/部件、锤钉、使用电动螺丝刀、休息手臂、转动螺丝刀和使用扳手。在该数据集上对所开发的多模态方法进行了评估,在留一和半半实验中,识别准确率分别高达97%和100%。
This study aims at sensing and understanding the worker’s activity in a human-centered intelligent manufacturing system. We propose a novel multi-modal approach for worker activity recognition by leveraging information from different sensors and in different modalities. Specifically, a smart armband and a visual camera are applied to capture Inertial Measurement Unit (IMU) signals and videos, respectively. For the IMU signals, we design two novel feature transform mechanisms, in both frequency and spatial domains, to assemble the captured IMU signals as images, which allow using convolutional neural networks to learn the most discriminative features. Along with the above two modalities, we propose two other modalities for the video data, i.e., at the video frame and video clip levels. Each of the four modalities returns a probability distribution on activity prediction. Then, these probability distributions are fused to output the worker activity classification result. A worker activity dataset is established, which at present contains 6 common activities in assembly tasks, i.e., grab a tool/part, hammer a nail, use a power-screwdriver, rest arms, turn a screwdriver, and use a wrench. The developed multi-modal approach is evaluated on this dataset and achieves recognition accuracies as high as 97% and 100% in the leave-one-out and half-half experiments, respectively.