Holistically-Attracted Wireframe Parsing: From Supervised to Self-Supervised Learning

Holistically-Attracted Wireframe Parsing: From Supervised to Self-Supervised Learning
复制标题

DOI:
10.1109/tpami.2023.3312749
复制
发表时间:
2022-10
影响因子:
23.6
通讯作者:
Nan Xue;Tianfu Wu;Song Bai;Fu-Dong Wang;Gui-Song Xia;L. Zhang;Philip H. S. Torr
Nan Xue;Tianfu Wu;Song Bai;Fu-Dong Wang;Gui-Song Xia;L. Zhang;Philip H. S. Torr
中科院分区:
计算机科学1区
文献类型:
--
作者:
Nan Xue;Tianfu Wu;Song Bai;Fu-Dong Wang;Gui-Song Xia;L. Zhang;Philip H. S. Torr

文献摘要

相似文献

提出了一种基于整体吸引线框分析(HAWP)的二维图像几何分析方法,该方法包含由线段和结点组成的线框。HAWP利用一种简洁的整体吸引(HAT)场表示法,它使用闭合形式的4D几何矢量场来编码线段。HAWP由端到端和HAT驱动的设计支持的三个顺序组件组成:1)从HAT场和热图的端点建议生成密集线段集,2)将密集线段绑定到稀疏端点建议以生成初始线框,以及3)通过一种新颖的端点去耦合兴趣线对齐(EPD LOIAlign)模块过滤虚假阳性建议,该模块捕获端点建议和HAT字段之间的共现以进行更好的验证。由于我们新颖的设计,HAWPv2在全监督学习中表现出强大的性能,而HAWPv3在自我监督学习方面表现出色,实现了卓越的可重复性分数和高效的训练(单个GPU上24小时的GPU)。此外,HAWPv3在不提供线框的地面真值标签的情况下,显示出在非分布图像中进行线框解析的潜力。
This article presents Holistically-Attracted Wireframe Parsing (HAWP), a method for geometric analysis of 2D images containing wireframes formed by line segments and junctions. HAWP utilizes a parsimonious Holistic Attraction (HAT) field representation that encodes line segments using a closed-form 4D geometric vector field. The proposed HAWP consists of three sequential components empowered by end-to-end and HAT-driven designs: 1) generating a dense set of line segments from HAT fields and endpoint proposals from heatmaps, 2) binding the dense line segments to sparse endpoint proposals to produce initial wireframes, and 3) filtering false positive proposals through a novel endpoint-decoupled line-of-interest aligning (EPD LOIAlign) module that captures the co-occurrence between endpoint proposals and HAT fields for better verification. Thanks to our novel designs, HAWPv2 shows strong performance in fully supervised learning, while HAWPv3 excels in self-supervised learning, achieving superior repeatability scores and efficient training (24 GPU hours on a single GPU). Furthermore, HAWPv3 exhibits a promising potential for wireframe parsing in out-of-distribution images without providing ground truth labels of wireframes.