Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges

Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges
复制标题

航空图像中的目标检测:大规模基准和挑战

DOI:
10.1109/tpami.2021.3117983
复制
发表时间:
2022-11-01
影响因子:
23.6
通讯作者:
Zhang, Liangpei
Zhang, Liangpei
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ding, Jian;Xue, Nan;Zhang, Liangpei

文献摘要

被引文献

相似文献

在过去的十年中,目标检测在自然图像中取得了重大进展,但在航空图像中却没有取得进展,这是由于航空图像的鸟瞰图导致物体的尺度和方向发生了巨大变化。更重要的是,缺乏大规模基准已经成为航空图像中目标检测(ODAI)发展的主要障碍。在本文中,我们提出了一个大规模的航空图像目标检测数据集(DOTA)和ODAI的综合基线。提出的DOTA数据集包含从11,268张航空图像中收集的18类定向边界框注释的1,793,658个对象实例。基于此大规模且注释良好的数据集,我们构建了涵盖70多种配置的10种最先进算法的基线,其中评估了每个模型的速度和准确性性能。此外,我们还为ODAI提供了一个代码库,并建立了一个评估不同算法的网站。DOTA以往的挑战已经吸引了全球1300多支队伍。我们认为,扩展的大规模DOTA数据集、广泛的基线、代码库和挑战可以促进航空图像中目标检测问题的鲁棒算法设计和可重复性研究。
In he past decade, object detection has achieved significant progress in natural images but not in aerial images, due to the massive variations in the scale and orientation of objects caused by the bird's-eye view of aerial images. More importantly, the lack of large-scale benchmarks has become a major obstacle to the development of object detection in aerial images (ODAI). In this paper, we present a large-scale Dataset of Object deTection in Aerial images (DOTA) and comprehensive baselines for ODAI. The proposed DOTA dataset contains 1,793,658 object instances of 18 categories of oriented-bounding-box annotations collected from 11,268 aerial images. Based on this large-scale and well-annotated dataset, we build baselines covering 10 state-of-the-art algorithms with over 70 configurations, where the speed and accuracy performances of each model have been evaluated. Furthermore, we provide a code library for ODAI and build a website for evaluating different algorithms. Previous challenges run on DOTA have attracted more than 1300 teams worldwide. We believe that the expanded large-scale DOTA dataset, the extensive baselines, the code library and the challenges can facilitate the designs of robust algorithms and reproducible research on the problem of object detection in aerial images.