Object Detection for Comics using Manga109 Annotations

Object Detection for Comics using Manga109 Annotations
复制标题

DOI:
--
复制
发表时间:
2018-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Toru Ogawa;Atsushi Otsubo;Rei Narita;Yusuke Matsui;T. Yamasaki;K. Aizawa
Toru Ogawa;Atsushi Otsubo;Rei Narita;Yusuke Matsui;T. Yamasaki;K. Aizawa
中科院分区:
其他
文献类型:
--
作者:
Toru Ogawa;Atsushi Otsubo;Rei Narita;Yusuke Matsui;T. Yamasaki;K. Aizawa

文献摘要

被引文献

相似文献

随着数字化漫画的发展,图像理解技术变得越来越重要。在本文中,我们专注于目标检测,这是一个基本的任务,图像理解。虽然基于卷积神经网络(CNN)的方法在自然图像的对象检测中表现良好,但将这些方法应用于漫画对象检测任务存在两个问题。首先,没有大规模的注释漫画数据集。基于CNN的方法需要大规模的注释进行训练。其次,漫画中的物体与自然主义图像相比高度重叠。这种重叠导致现有的基于CNN的方法中的分配问题。为了解决这些问题,我们提出了一个新的注释数据集和一个新的CNN模型。我们对现有的漫画图像数据集进行了注释,并创建了最大的注释数据集,名为Manga 109-annotations。对于分配问题,我们提出了一种新的基于CNN的检测器,SSD 300-fork。我们将SSD 300-fork与使用Manga 109-annotations的其他检测方法进行了比较,并确认我们的模型基于mAP评分优于它们。
With the growth of digitized comics, image understanding techniques are becoming important. In this paper, we focus on object detection, which is a fundamental task of image understanding. Although convolutional neural networks (CNN)-based methods archived good performance in object detection for naturalistic images, there are two problems in applying these methods to the comic object detection task. First, there is no large-scale annotated comics dataset. The CNN-based methods require large-scale annotations for training. Secondly, the objects in comics are highly overlapped compared to naturalistic images. This overlap causes the assignment problem in the existing CNN-based methods. To solve these problems, we proposed a new annotation dataset and a new CNN model. We annotated an existing image dataset of comics and created the largest annotation dataset, named Manga109-annotations. For the assignment problem, we proposed a new CNN-based detector, SSD300-fork. We compared SSD300-fork with other detection methods using Manga109-annotations and confirmed that our model outperformed them based on the mAP score.