Semantic Understanding of Scenes Through the ADE20K Dataset

Semantic Understanding of Scenes Through the ADE20K Dataset
复制标题

DOI:
10.1007/s11263-018-1140-0
复制
发表时间:
2019-03-01
影响因子:
19.5
通讯作者:
Torralba, Antonio
Torralba, Antonio
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhou, Bolei;Zhao, Hang;Torralba, Antonio

文献摘要

被引文献

相似文献

视觉场景的语义理解是计算机视觉的圣杯之一。尽管社区在数据收集方面做出了努力,但仍然很少有涵盖广泛场景和对象类别的图像数据集,这些图像数据具有像素级注释以用于场景理解。在这项工作中,我们提出了一个密集标注的数据集ADE20K,它跨越了场景、对象、对象的部分,在某些情况下甚至部分的标注。总共有25K张复杂的日常场景的图像,其中包含自然空间环境中的各种对象。平均每幅图像有19.5个实例和10.5个对象类。基于ADE20K构建了场景解析和实例分割的基准测试程序。我们提供了这两个基准测试的基准性能,并为开源重新实现了最先进的模型。我们进一步评估了同步批归一化的效果,发现合理的大批大小是语义分割性能的关键。实验结果表明,在ADE20K上训练的网络能够分割出各种场景和对象。
Semantic understanding of visual scenes is one of the holy grails of computer vision. Despite efforts of the community in data collection, there are still few image datasets covering a wide range of scenes and object categories with pixel-wise annotations for scene understanding. In this work, we present a densely annotated dataset ADE20K, which spans diverse annotations of scenes, objects, parts of objects, and in some cases even parts of parts. Totally there are 25k images of the complex everyday scenes containing a variety of objects in their natural spatial context. On average there are 19.5 instances and 10.5 object classes per image. Based on ADE20K, we construct benchmarks for scene parsing and instance segmentation. We provide baseline performances on both of the benchmarks and re-implement state-of-the-art models for open source. We further evaluate the effect of synchronized batch normalization and find that a reasonably large batch size is crucial for the semantic segmentation performance. We show that the networks trained on ADE20K are able to segment a wide variety of scenes and objects.