Shared Manifold Learning Using a Triplet Network for Multiple Sensor Translation and Fusion With Missing Data

Shared Manifold Learning Using a Triplet Network for Multiple Sensor Translation and Fusion With Missing Data
复制标题

DOI:
10.1109/jstars.2022.3217485
复制
发表时间:
2022-10
影响因子:
5.5
通讯作者:
Aditya Dutt;Alina Zare;P. Gader
Aditya Dutt;Alina Zare;P. Gader
中科院分区:
工程技术3区
文献类型:
--
作者:
Aditya Dutt;Alina Zare;P. Gader

文献摘要

相似文献

异构数据融合可以提高算法在特定任务上的鲁棒性和准确性。然而,由于各种模态的差异,对齐传感器并将其信息嵌入到有区别的紧凑表示中是具有挑战性的。在这篇文章中,我们提出了一个基于对比学习的多模态对齐网络,将来自不同传感器的数据对齐到一个共享的和有区别的流形中,其中保留了类别信息。所提出的架构使用多模态三重自动编码器集群的潜在空间,在这样一种方式,从每个异构模态的相同类的样本被映射到彼此接近。由于所有的模态存在于一个共享的流形,提出了一个统一的分类框架。将得到的潜在空间表示进行融合,以执行更鲁棒和准确的分类。在缺失传感器的情况下,使用另一传感器的潜在空间容易且有效地预测一个传感器的潜在空间,从而允许传感器平移。我们进行了大量的实验,手动标记的多模态数据集包含高光谱数据从AVIRIS-NG和氖和光探测和测距(LiDAR)数据从氖。最后,该模型在两个基准数据集上进行了验证:柏林数据集(高光谱和合成孔径雷达)和MUUFL格尔夫波特数据集(高光谱和激光雷达)。通过与其它方法的比较,证明了该方法的优越性。我们在MUUFL数据集上实现了94.3%的平均总体准确度,在柏林数据集上实现了71.26%的最佳总体准确度,这优于其他最先进的方法。
Heterogeneous data fusion can enhance the robustness and accuracy of an algorithm on a given task. However, due to the difference in various modalities, aligning the sensors and embedding their information into discriminative and compact representations is challenging. In this article, we propose a contrastive learning-based multimodal alignment network to align data from different sensors into a shared and discriminative manifold where class information is preserved. The proposed architecture uses a multimodal triplet autoencoder to cluster the latent space in such a way that samples of the same classes from each heterogeneous modality are mapped close to each other. Since all the modalities exist in a shared manifold, a unified classification framework is proposed. The resulting latent space representations are fused to perform more robust and accurate classification. In a missing sensor scenario, the latent space of one sensor is easily and efficiently predicted using another sensor's latent space, thereby allowing sensor translation. We conducted extensive experiments on a manually labeled multimodal dataset containing hyperspectral data from AVIRIS-NG and NEON and light detection and ranging (LiDAR) data from NEON. Finally, the model is validated on two benchmark datasets: Berlin Dataset (hyperspectral and synthetic aperture radar) and MUUFL Gulfport Dataset (hyperspectral and LiDAR). A comparison made with other methods demonstrates the superiority of this method. We achieved a mean overall accuracy of 94.3% on the MUUFL dataset and the best overall accuracy of 71.26% on the Berlin dataset, which is better than other state-of-the-art approaches.