The PASCAL Visual Object Classes Challenge: A Retrospective

The PASCAL Visual Object Classes Challenge: A Retrospective
复制标题

DOI:
10.1007/s11263-014-0733-5
复制
发表时间:
2015-01-01
影响因子:
19.5
通讯作者:
Zisserman, Andrew
Zisserman, Andrew
中科院分区:
计算机科学2区
文献类型:
--
作者:
Everingham, Mark;Eslami, S. M. Ali;Zisserman, Andrew

文献摘要

被引文献

相似文献

Pascal可视化对象类(VOC)挑战由两个部分组成:(i)公开可用的图像数据集,以及地面真相注释和标准化评估软件;(ii)年度竞赛和研讨会。有五个挑战:分类、检测、分割、动作分类和人员布局。在本文中,我们对2008-2012年的挑战进行了回顾。本文主要针对两类读者:算法设计者,希望了解当前技术现状的研究人员,通过VOC数据集的性能来衡量,以及当前一代算法的局限性和弱点;还有挑战设计师,他们想看看我们作为组织者从这个过程中学到了什么,以及我们对未来挑战组织的建议。为了分析提交的算法在VOC数据集上的性能,我们引入了一些新的评估方法:一种用于确定两种算法性能差异是否显著的自举方法;标准化的平均精度,以便可以在具有不同比例的正实例的类别之间比较性能;一种用于跨多个算法可视化性能的聚类方法,以便可以识别困难和容易的图像;并对提交的算法使用联合分类器来衡量它们的互补性和综合性能。我们还使用Hoiem等人(欧洲计算机视觉会议论文集,2012)的方法分析了该社区的进展情况,以确定发生错误的类型。最后,我们对挑战中运作良好的方面以及在未来挑战中可以改进的方面进行了评估。
The Pascal Visual Object Classes (VOC) challenge consists of two components: (i) a publicly available dataset of images together with ground truth annotation and standardised evaluation software; and (ii) an annual competition and workshop. There are five challenges: classification, detection, segmentation, action classification, and person layout. In this paper we provide a review of the challenge from 2008-2012. The paper is intended for two audiences: algorithm designers, researchers who want to see what the state of the art is, as measured by performance on the VOC datasets, along with the limitations and weak points of the current generation of algorithms; and, challenge designers, who want to see what we as organisers have learnt from the process and our recommendations for the organisation of future challenges. To analyse the performance of submitted algorithms on the VOC datasets we introduce a number of novel evaluation methods: a bootstrapping method for determining whether differences in the performance of two algorithms are significant or not; a normalised average precision so that performance can be compared across classes with different proportions of positive instances; a clustering method for visualising the performance across multiple algorithms so that the hard and easy images can be identified; and the use of a joint classifier over the submitted algorithms in order to measure their complementarity and combined performance. We also analyse the community's progress through time using the methods of Hoiem et al. (Proceedings of European Conference on Computer Vision, 2012) to identify the types of occurring errors. We conclude the paper with an appraisal of the aspects of the challenge that worked well, and those that could be improved in future challenges.