E3Pose: Energy-Efficient Edge-assisted Multi-camera System for Multi-human 3D Pose Estimation

E3Pose: Energy-Efficient Edge-assisted Multi-camera System for Multi-human 3D Pose Estimation
复制标题

DOI:
10.1145/3576842.3582370
复制
发表时间:
2023-01
期刊:
Proceedings of the 8th ACM/IEEE Conference on Internet of Things Design and Implementation
影响因子:
--
通讯作者:
Letian Zhang;Jie Xu
Letian Zhang;Jie Xu
中科院分区:
其他
文献类型:
--
作者:
Letian Zhang;Jie Xu

文献摘要

相似文献

多人三维位姿估计对于建立真实的世界和虚拟世界之间的无缝连接起着关键作用。最近的努力采用了一个两阶段的框架,首先从不同的角度在多个相机视图中建立2D姿态估计,然后将它们合成为3D姿态。然而,重点主要是在离线视频数据集上开发新的计算机视觉算法,而没有过多考虑具有灵活部署和电池供电相机的现实系统中的能源约束。在本文中,我们提出了一个节能的边缘辅助多摄像头系统,被称为E3Pose,实时多人3D姿态估计,基于自适应相机选择的关键思想。E3Pose并不像现有技术那样总是使用所有可用的摄像机来执行2D姿态估计,而是根据摄像机的遮挡和能量状态,以自适应的方式选择摄像机的子集,从而降低能耗(这意味着延长了电池寿命)并提高估计精度。为了实现这一目标,E3Pose采用了基于注意力的LSTM来预测每个摄像机视图的遮挡信息,并在选择摄像机处理场景图像之前指导摄像机选择,并运行基于李亚普诺夫优化框架的摄像机选择算法,以做出长期自适应选择决策。我们建立了一个原型的E3Pose的5摄像头测试平台上,证明了它的可行性,并评估其性能。我们的研究结果表明,可以实现显着的节能(高达31.21%),同时保持高的3D姿态估计精度相媲美的国家的最先进的方法。
Multi-human 3D pose estimation plays a key role in establishing a seamless connection between the real world and the virtual world. Recent efforts adopted a two-stage framework that first builds 2D pose estimations in multiple camera views from different perspectives and then synthesizes them into 3D poses. However, the focus has largely been on developing new computer vision algorithms on the offline video datasets without much consideration on the energy constraints in real-world systems with flexibly-deployed and battery-powered cameras. In this paper, we propose an energy-efficient edge-assisted multiple-camera system, dubbed E3Pose, for real-time multi-human 3D pose estimation, based on the key idea of adaptive camera selection. Instead of always employing all available cameras to perform 2D pose estimations as in the existing works, E3Pose selects only a subset of cameras depending on their camera view qualities in terms of occlusion and energy states in an adaptive manner, thereby reducing the energy consumption (which translates to extended battery lifetime) and improving the estimation accuracy. To achieve this goal, E3Pose incorporates an attention-based LSTM to predict the occlusion information of each camera view and guide camera selection before cameras are selected to process the images of a scene, and runs a camera selection algorithm based on the Lyapunov optimization framework to make long-term adaptive selection decisions. We build a prototype of E3Pose on a 5-camera testbed, demonstrate its feasibility and evaluate its performance. Our results show that a significant energy saving (up to 31.21%) can be achieved while maintaining a high 3D pose estimation accuracy comparable to state-of-the-art methods.