Vector Quantization to Visualize the Detection Process

Vector Quantization to Visualize the Detection Process
复制标题

DOI:
10.5220/0010426005530561
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Kaiyu Suzuki;Tomofumi Matsuzawa;M. Takimoto;Y. Kambayashi
Kaiyu Suzuki;Tomofumi Matsuzawa;M. Takimoto;Y. Kambayashi
中科院分区:
其他
文献类型:
--
作者:
Kaiyu Suzuki;Tomofumi Matsuzawa;M. Takimoto;Y. Kambayashi

文献摘要

相似文献

无人机最重要的任务之一是根据机载摄像头捕获的图像自动做出决策,并为疏散人员提供有用的信息,例如疏散指导。为了从上述图像中自动做出决策,深度学习是最合适和最强大的方法。虽然深度学习表现出很高的性能,但提出决策的理由是一个挑战。尽管现有的几种决策方法可视化并指出他们集中考虑了图像的哪一部分,但对于需要紧急和准确判断的情况来说,它们是不够的。当我们寻找决策的依据时,我们不仅需要知道在哪里检测,还需要知道如何检测。本研究旨在将矢量量化(VQ)插入中间层作为第一步,以展示如何在基于图像的任务中检测深度学习。我们提出了一种方法,抑制精度损失,同时保持可解释性,通过应用VQ的分类问题。Sinkhorn-Knopp算法,常数嵌入空间和梯度惩罚在这项研究中的应用,使我们能够引入VQ与高解释性。这些技术应该可以帮助我们将所提出的方法应用于现实世界的任务,其中数据集的属性是
One of the most important tasks for drones, which are in the spotlight for assisting evacuees of natural disasters, is to automatically make decisions based on images captured by on-board cameras and provide evacuees with useful information, such as evacuation guidance. In order to make decision automatically from the aforementioned images, deep learning is the most suitable and powerful method. Although deep learning exhibits high performance, presenting the rationale for decisions is a challenge. Even though several existing decision making methods visualize and point out which part of the image they have considered intensively, they are insufficient for situations that require urgent and accurate judgments. When we look for basis for the decisions, we need to know not only WHERE to detect but also HOW to detect. This study aims to insert vector quantization (VQ) into the intermediate layer as a first step in order to show HOW to detect for deep learning in image-based tasks. We propose a method that suppresses accuracy loss while holding interpretability by applying VQ to the classification problem. The applications of the Sinkhorn–Knopp algorithm, constant embedding space and gradient penalty in this study allow us to introduce VQ with high interpretability. These techniques should help us apply the proposed method to real-world tasks where the properties of datasets are