GeoCLR: Georeference Contrastive Learning for Efficient Seafloor Image Interpretation

GeoCLR: Georeference Contrastive Learning for Efficient Seafloor Image Interpretation
复制标题

DOI:
10.55417/fr.2022037
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Takaki Yamada;A. Prügel-Bennett;Stefan B. Williams;O. Pizarro;B. Thornton
Takaki Yamada;A. Prügel-Bennett;Stefan B. Williams;O. Pizarro;B. Thornton
中科院分区:
其他
文献类型:
--
作者:
Takaki Yamada;A. Prügel-Bennett;Stefan B. Williams;O. Pizarro;B. Thornton

文献摘要

相似文献

本文描述了用于深度学习卷积神经网络(CNN)的高效训练的视觉表示的地理参考对比学习(Geoferencing)。该方法通过使用在附近位置拍摄的图像生成相似的图像对,并将这些图像与相距较远的图像对进行对比,来利用地理参考信息。基本的假设是,在近距离内收集的图像更有可能具有相似的视觉外观,这在海底机器人成像应用中可以合理地得到满足,其中图像足迹被限制在几米的边缘长度,并且被采取,以便它们沿着车辆的轨迹重叠,而海底基质和栖息地的斑块尺寸要大得多。这种方法的一个关键优点是它是自我监督的,不需要任何人工输入来进行CNN训练。该方法是计算效率高,结果可以产生在多日的自主水下航行器(AUV)任务使用的计算资源,将在大多数海洋现场试验期间访问的潜水之间。我们将Geoprock应用于由AUV收集的约86,000张图像组成的数据集的栖息地分类。我们演示了如何使用Geoconstruct生成的潜在表示来有效地指导人类注释工作,其中半监督框架与使用相同CNN和相同数量的人类注释进行训练的最先进Simonstruct相比,平均提高了10.2%的分类准确性。
This paper describes georeference contrastive learning of visual representation (GeoCLR) for efficient training of deep-learning convolutional neural networks (CNNs). The method leverages georeference information by generating a similar image pair using images taken of nearby locations, and contrasting these with an image pair that is far apart. The underlying assumption is that images gathered within a close distance are more likely to have similar visual appearance, where this can be reasonably satisfied in seafloor robotic imaging applications where image footprints are limited to edge lengths of a few meters and are taken so that they overlap along a vehicle’s trajectory, whereas seafloor substrates and habitats have patch sizes that are far larger. A key advantage of this method is that it is self-supervised and does not require any human input for CNN training. The method is computationally efficient, where results can be generated between dives during multi-day autonomous underwater vehicle (AUV) missions using computational resources that would be accessible during most oceanic field trials. We apply GeoCLR to habitat classification on a dataset that consists of ~86,000 images gathered using an AUV. We demonstrate how the latent representations generated by GeoCLR can be used to efficiently guide human annotation efforts, where the semi-supervised framework improves classification accuracy by an average of 10.2% compared to the state-of-the-art SimCLR using the same CNN and equivalent number of human annotations for training.