Instance Segmentation, Body Part Parsing, and Pose Estimation of Human Figures in Pictorial Maps

Instance Segmentation, Body Part Parsing, and Pose Estimation of Human Figures in Pictorial Maps
复制标题

DOI:
10.1080/23729333.2021.1949087
复制
发表时间:
2021-08
影响因子:
0.5
通讯作者:
R. Schnürer;A. Cengiz Öztireli;M. Heitzler;R. Sieber;L. Hurni
R. Schnürer;A. Cengiz Öztireli;M. Heitzler;R. Sieber;L. Hurni
中科院分区:
--
文献类型:
--
作者:
R. Schnürer;A. Cengiz Öztireli;M. Heitzler;R. Sieber;L. Hurni

文献摘要

相似文献

摘要 近年来,卷积神经网络(CNN)已成功应用于识别人物及其身体部位以及照片和视频中的关键点。将这些技术转移到人工创建的图像上还没有被探索过,尽管具有挑战性,因为这些图像是以不同的风格、身体比例和抽象层次绘制的。在这项工作中,我们在图像地图的基础上研究这些问题,其中我们使用两个连续的 CNN 识别包含的人物:我们首先使用 Mask R-CNN 分割单个人物,然后解析他们的身体部位并使用四个不同的 UNet++ 版本同时估计他们的姿势。我们使用真人和合成人物的混合来训练 CNN,并将结果与​​由图形组成的手动注释的测试数据集进行比较。通过改变训练数据集和 CNN 配置,我们能够改进原始的 Mask R-CNN 模型,并且我们在 UNet++ 版本中取得了令人满意的结果。提取的图形可用于动画和讲故事,并且可能与历史和当代地图的分析相关。
ABSTRACT In recent years, convolutional neural networks (CNNs) have been applied successfully to recognise persons, their body parts and pose keypoints in photos and videos. The transfer of these techniques to artificially created images is rather unexplored, though challenging since these images are drawn in different styles, body proportions, and levels of abstraction. In this work, we study these problems on the basis of pictorial maps where we identify included human figures with two consecutive CNNs: We first segment individual figures with Mask R-CNN, and then parse their body parts and estimate their poses simultaneously with four different UNet++ versions. We train the CNNs with a mixture of real persons and synthetic figures and compare the results with manually annotated test datasets consisting of pictorial figures. By varying the training datasets and the CNN configurations, we were able to improve the original Mask R-CNN model and we achieved moderately satisfying results with the UNet++ versions. The extracted figures may be used for animation and storytelling and may be relevant for the analysis of historic and contemporary maps.