Weakly Supervised Object Localization with Multi-Fold Multiple Instance Learning

Weakly Supervised Object Localization with Multi-Fold Multiple Instance Learning
复制标题

DOI:
10.1109/tpami.2016.2535231
复制
发表时间:
2017-01-01
影响因子:
23.6
通讯作者:
Schmid, Cordelia
Schmid, Cordelia
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cinbis, Ramazan Gokberk;Verbeek, Jakob;Schmid, Cordelia

文献摘要

被引文献

相似文献

目标类别定位是计算机视觉中的一个具有挑战性的问题。标准的监督训练需要对象实例的边界框注释。在弱监督学习中,这种耗时的标注过程被绕过。在这种情况下,监督信息被限制为指示图像中对象实例的不存在/存在的二进制标签,而不包括它们的位置。我们遵循多示例学习的方法,迭代地训练检测器,并在正向训练图像中推断目标位置。我们的主要贡献是一种多重多实例学习过程,它防止训练过早锁定错误的对象位置。当使用诸如Fisher矢量和卷积神经网络特征之类的高维表示时,该过程尤其重要。我们还提出了一种窗口细化方法,通过引入客观性先验来提高定位精度。我们使用PASCALVOC 2007数据集进行了详细的实验评估,验证了该方法的有效性。
Object category localization is a challenging problem in computer vision. Standard supervised training requires bounding box annotations of object instances. This time-consuming annotation process is sidestepped in weakly supervised learning. In this case, the supervised information is restricted to binary labels that indicate the absence/presence of object instances in the image, without their locations. We follow a multiple-instance learning approach that iteratively trains the detector and infers the object locations in the positive training images. Our main contribution is a multi-fold multiple instance learning procedure, which prevents training from prematurely locking onto erroneous object locations. This procedure is particularly important when using high-dimensional representations, such as Fisher vectors and convolutional neural network features. We also propose a window refinement method, which improves the localization accuracy by incorporating an objectness prior. We present a detailed experimental evaluation using the PASCALVOC 2007 dataset, which verifies the effectiveness of our approach.