Zero-shot Building Attribute Extraction from Large-Scale Vision and Language Models

Zero-shot Building Attribute Extraction from Large-Scale Vision and Language Models
复制标题

DOI:
10.1109/wacv57701.2024.00845
复制
发表时间:
2023-12
期刊:
2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Fei Pan;Sangryul Jeon;Brian Wang;Frank Mckenna;Stella X. Yu
Fei Pan;Sangryul Jeon;Brian Wang;Frank Mckenna;Stella X. Yu
中科院分区:
其他
文献类型:
--
作者:
Fei Pan;Sangryul Jeon;Brian Wang;Frank Mckenna;Stella X. Yu

文献摘要

相似文献

现有的建筑物识别方法,以BRAILS为例,利用监督学习从卫星和街景图像中提取信息进行分类和分割。然而,每个任务模块都需要人工注释的数据,这阻碍了对区域差异和注释不平衡的可扩展性和鲁棒性。作为回应,我们提出了一种新的零采样工作流来构建属性提取,该工作流利用大规模的视觉和语言模型来减轻对外部注释的依赖。该工作流包含两个关键部分:基于结构和土木工程相关词汇表的图像级字幕和片段级字幕。这两个组件通过计算图像和词汇表的特征表示来生成描述性标题,并促进视觉表示和文本表示之间的语义匹配。因此,我们的框架为结构和土木工程领域的建筑属性提取提供了一个有前途的途径,最终减少了对人类注释的依赖,同时提高了性能和适应性。
Existing building recognition methods, exemplified by BRAILS, utilize supervised learning to extract information from satellite and street-view images for classification and segmentation. However, each task module requires human-annotated data, hindering the scalability and robustness to regional variations and annotation imbalances. In response, we propose a new zero-shot workflow for building attribute extraction that utilizes large-scale vision and language models to mitigate reliance on external annotations. The proposed workflow contains two key components: image-level captioning and segment-level captioning for the building images based on the vocabularies pertinent to structural and civil engineering. These two components generate descriptive captions by computing feature representations of the image and the vocabularies, and facilitating a semantic match between the visual and textual representations. Consequently, our framework offers a promising avenue to enhance AI-driven captioning for building attribute extraction in the structural and civil engineering domains, ultimately reducing reliance on human annotations while bolstering performance and adaptability.