Improving Deep Learning for Maritime Remote Sensing through Data Augmentation and Latent Space

Improving Deep Learning for Maritime Remote Sensing through Data Augmentation and Latent Space
复制标题

DOI:
10.3390/make4030031
复制
发表时间:
2022-07
期刊:
Mach. Learn. Knowl. Extr.
影响因子:
--
通讯作者:
Daniel Sobien;Erik Higgins;J. Krometis;Justin Kauffman;Laura J. Freeman
Daniel Sobien;Erik Higgins;J. Krometis;Justin Kauffman;Laura J. Freeman
中科院分区:
其他
文献类型:
--
作者:
Daniel Sobien;Erik Higgins;J. Krometis;Justin Kauffman;Laura J. Freeman

文献摘要

被引文献

相似文献

训练深度学习模型需要有正确的问题数据,并理解你的数据和模型在数据上的表现。当数据有限时,训练深度学习模型是困难的,因此在本文中,我们寻求回答以下问题:如何在有限的数据下训练深度学习模型以提高其在目标区域的性能?我们通过对模拟合成孔径雷达(SAR)图像数据集应用旋转数据增强来做到这一点。我们使用统一流形逼近和投影(UMAP)降维技术来理解增广对潜在空间数据的影响。使用这种潜在空间表示,我们可以理解数据并选择特定的训练样本,目的是在目标表现不佳的区域提高模型的性能,而无需增加训练集的大小。结果表明,在某些情况下,使用潜在空间选择训练数据可以显著提高模型的性能;然而,在其他情况下,没有任何改进。我们表明,潜在空间中的链接模式是模型性能的可能预测因子,但结果需要一些实验和领域知识来确定最佳选择。
Training deep learning models requires having the right data for the problem and understanding both your data and the models’ performance on that data. Training deep learning models is difficult when data are limited, so in this paper, we seek to answer the following question: how can we train a deep learning model to increase its performance on a targeted area with limited data? We do this by applying rotation data augmentations to a simulated synthetic aperture radar (SAR) image dataset. We use the Uniform Manifold Approximation and Projection (UMAP) dimensionality reduction technique to understand the effects of augmentations on the data in latent space. Using this latent space representation, we can understand the data and choose specific training samples aimed at boosting model performance in targeted under-performing regions without the need to increase training set sizes. Results show that using latent space to choose training data significantly improves model performance in some cases; however, there are other cases where no improvements are made. We show that linking patterns in latent space is a possible predictor of model performance, but results require some experimentation and domain knowledge to determine the best options.