DBAT: Dynamic Backward Attention Transformer for Material Segmentation with Cross-Resolution Patches

DBAT: Dynamic Backward Attention Transformer for Material Segmentation with Cross-Resolution Patches
复制标题

DOI:
10.48550/arxiv.2305.03919
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuwen Heng;S. Dasmahapatra;Hansung Kim
Yuwen Heng;S. Dasmahapatra;Hansung Kim
中科院分区:
其他
文献类型:
--
作者:
Yuwen Heng;S. Dasmahapatra;Hansung Kim

文献摘要

相似文献

密集材料分割的目标是识别每个图像像素的材料类别。最近的研究采用图像贴片提取材料特征。虽然训练的网络可以提高分割性能,但它们的方法选择了固定的patch分辨率,没有考虑到每种材料所覆盖的像素面积的变化。在本文中,我们提出了动态后向注意转换器(DBAT)来聚合交叉分辨率特征。该算法以裁剪后的图像块为输入,在每个变换阶段通过合并相邻的块逐渐提高块的分辨率,而不是在训练时固定块的分辨率。我们明确地收集从交叉分辨率补丁中提取的中间特征,并动态地将它们与预测的注意力掩模合并。实验表明,该算法的准确率达到了86.85%,是目前最先进的实时模型中准确率最高的。与其他具有复杂架构的成功深度学习解决方案一样,DBAT也存在缺乏可解释性的问题。为了解决这个问题,本文研究了DBAT使用的属性。本文通过分析交叉分辨率特征和注意权值,解释了DBAT如何从图像块中学习。我们进一步将特征与语义标签对齐,进行网络解剖,以推断所提出的模型可以比其他方法更好地提取与材料相关的特征。我们表明,与其他模型相比,DBAT模型对网络初始化更具鲁棒性,并且产生更少的变量预测。项目代码可从https://github.com/heng-yuwen/Dynamic-Backward-Attention-Transformer获得。
The objective of dense material segmentation is to identify the material categories for every image pixel. Recent studies adopt image patches to extract material features. Although the trained networks can improve the segmentation performance, their methods choose a fixed patch resolution which fails to take into account the variation in pixel area covered by each material. In this paper, we propose the Dynamic Backward Attention Transformer (DBAT) to aggregate cross-resolution features. The DBAT takes cropped image patches as input and gradually increases the patch resolution by merging adjacent patches at each transformer stage, instead of fixing the patch resolution during training. We explicitly gather the intermediate features extracted from cross-resolution patches and merge them dynamically with predicted attention masks. Experiments show that our DBAT achieves an accuracy of 86.85%, which is the best performance among state-of-the-art real-time models. Like other successful deep learning solutions with complex architectures, the DBAT also suffers from lack of interpretability. To address this problem, this paper examines the properties that the DBAT makes use of. By analysing the cross-resolution features and the attention weights, this paper interprets how the DBAT learns from image patches. We further align features to semantic labels, performing network dissection, to infer that the proposed model can extract material-related features better than other methods. We show that the DBAT model is more robust to network initialisation, and yields fewer variable predictions compared to other models. The project code is available at https://github.com/heng-yuwen/Dynamic-Backward-Attention-Transformer.