Structure-aware human pose estimation with graph convolutional networks

Structure-aware human pose estimation with graph convolutional networks
复制标题

使用图卷积网络进行结构感知人体姿势估计

DOI:
10.1016/j.patcog.2020.107410
复制
发表时间:
2020-10-01
影响因子:
8
通讯作者:
Sang, Nong
Sang, Nong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bin, Yanrui;Chen, Zhao-Min;Sang, Nong

文献摘要

被引文献

相似文献

人体姿态估计是从静止图像中定位人体关键点的任务。由于人体关键点是相互关联的,因此需要对人体关键点之间的结构关系进行建模,以进一步提高定位性能。在本文中,我们在原有的图卷积网络的基础上,提出了一种新的模型,称为位姿图卷积网络(PGCN),以利用这些重要的关系进行位姿估计。具体地说,我们的模型根据人体的自然组成模型建立了人体关键点之间的有向图。每个节点(关键点)由一个由多个特征地图组成的3-D张量表示,这些特征地图最初是由我们的骨干网络生成的,以保持准确的空间信息。此外,还提出了关注关键点之间的关键边(结构化信息)的注意机制。然后,学习PGCN将图形映射成一组结构感知关键点表示,这些关键点表示既编码了人体的结构,又编码了特定关键点的外观信息。此外,我们还提出了两个PGCN模块,即本地PGCN(L-PGCN)模块和非本地PGCN(NL-PGCN)模块。前者利用空间注意力来捕捉相邻关键点的局部区域之间的相关性,从而细化关键点的位置。后者通过异地操作捕捉远程关系,将具有挑战性的关键点关联起来。通过配备这两个模块,我们的PGCN可以进一步提高本地化性能。在单人和多人评估基准数据集上的实验表明,我们的方法始终优于竞争最先进的方法。(C)2020爱思唯尔有限公司。保留所有权利。
Human pose estimation is the task of localizing body key points from still images. As body key points are inter-connected, it is desirable to model the structural relationships between body key points to further improve the localization performance. In this paper, based on original graph convolutional networks, we propose a novel model, termed Pose Graph Convolutional Network (PGCN), to exploit these important relationships for pose estimation. Specifically, our model builds a directed graph between body key points according to the natural compositional model of a human body. Each node (key point) is represented by a 3-D tensor consisting of multiple feature maps, initially generated by our backbone network, to retain accurate spatial information. Furthermore, attention mechanism is presented to focus on crucial edges (structured information) between key points. PGCN is then learned to map the graph into a set of structure-aware key point representations which encode both structure of human body and appearance information of specific key points. Additionally, we propose two modules for PGCN, i.e., the Local PGCN (L-PGCN) module and Non-Local PGCN (NL-PGCN) module. The former utilizes spatial attention to capture the correlations between the local areas of adjacent key points to refine the location of key points. While the latter captures long-range relationships via non-local operation to associate the challenging key points. By equipping with these two modules, our PGCN can further improve localization performance. Experiments both on single- and multi-person estimation benchmark datasets show that our method consistently outperforms competing state-of-the-art methods. (C) 2020 Elsevier Ltd. All rights reserved.