Autonomous Human Activity Classification From Wearable Multi-Modal Sensors

Autonomous Human Activity Classification From Wearable Multi-Modal Sensors
复制标题

DOI:
10.1109/jsen.2019.2934678
复制
发表时间:
2019-12-01
影响因子:
4.3
通讯作者:
Velipasalar, Senem
Velipasalar, Senem
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Lu, Yantao;Velipasalar, Senem

文献摘要

被引文献

相似文献

基于惯性测量单元 (IMU) 数据或提供第三人称视角的静态摄像机的数据,针对人类活动分类进行了大量研究工作。使用可穿戴相机提供第一人称或以自我为中心的视图的工作相对较少,而将自我中心视频与 IMU 数据相结合的方法则更少。仅使用 IMU 数据限制了可检测活动的多样性和复杂性。例如,通过IMU数据可以检测坐姿活动,但无法确定受试者是否坐在椅子上或沙发上,或者受试者在哪里。为了执行细粒度的活动分类,并区分仅通过 IMU 数据无法区分的活动,我们提出了一种使用来自可穿戴相机和 IMU 的数据的自主且稳健的方法。与基于卷积神经网络的方法相比,我们建议采用胶囊网络从以自我为中心的视频数据中获取特征。此外,卷积长短期记忆框架还用于以自我为中心的视频和 IMU 数据,以捕获动作的时间方面。我们还提出了一种基于遗传算法的方法来自主、系统地设置各种网络参数,而不是使用手动设置。已经进行了 9 标签和 26 标签活动分类的实验,并且所提出的方法使用自主设置的网络参数,提供了非常有希望的结果,分别实现了 86.6% 和 77.2% 的总体准确率。与仅使用自我视觉数据和仅使用 IMU 数据相比,所提出的方法结合了两种模式,还提供了更高的准确性。
There has been significant amount of research work on human activity classification relying either on Inertial Measurement Unit (IMU) data or data from static cameras providing a third-person view. There has been relatively less work using wearable cameras, providing first-person or egocentric view, and even fewer approaches combining egocentric video with IMU data. Using only IMU data limits the variety and complexity of the activities that can be detected. For instance, the sitting activity can be detected by IMU data, but it cannot be determined whether the subject has sat on a chair or a sofa, or where the subject is. To perform fine-grained activity classification, and to distinguish between activities that cannot be differentiated by only IMU data, we present an autonomous and robust method using data from both wearable cameras and IMUs. In contrast to convolutional neural network-based approaches, we propose to employ capsule networks to obtain features from egocentric video data. Moreover, Convolutional Long Short Term Memory framework is employed both on egocentric videos and IMU data to capture the temporal aspect of actions. We also propose a genetic algorithm-based approach to autonomously and systematically set various network parameters, rather than using manual settings. Experiments have been conducted to perform 9- and 26-label activity classification, and the proposed method, using autonomously set network parameters, has provided very promising results, achieving overall accuracies of 86.6% and 77.2%, respectively. The proposed approach, combining both modalities, also provides increased accuracy compared to using only egovision data and only IMU data.