Going Beyond Nouns With Vision & Language Models Using Synthetic Data

Going Beyond Nouns With Vision & Language Models Using Synthetic Data
复制标题

DOI:
10.1109/iccv51070.2023.01844
复制
发表时间:
2023-03
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Paola Cascante-Bonilla;Khaled Shehada;James Smith;Sivan Doveh;Donghyun Kim-;Rameswar Panda;Gül Varol;A. Oliva;Vicente Ordonez;R. Feris;Leonid Karlinsky
Paola Cascante-Bonilla;Khaled Shehada;James Smith;Sivan Doveh;Donghyun Kim-;Rameswar Panda;Gül Varol;A. Oliva;Vicente Ordonez;R. Feris;Leonid Karlinsky
中科院分区:
其他
文献类型:
--
作者:
Paola Cascante-Bonilla;Khaled Shehada;James Smith;Sivan Doveh;Donghyun Kim-;Rameswar Panda;Gül Varol;A. Oliva;Vicente Ordonez;R. Feris;Leonid Karlinsky

文献摘要

相似文献

大规模预训练的视觉和语言(VL)模型在许多应用中表现出了卓越的性能,可以用零射击开放词汇推理(几乎任意)的自然语言提示取代一组固定的支持类。然而,最近的研究发现了这些模型的一个根本弱点。例如,他们难以理解“超越名词”的视觉语言概念(VLC),如非对象词(如属性、动作、关系、状态等)的含义,或难以进行组合推理,如理解句子中单词顺序的重要性。在这项工作中,我们研究了在多大程度上可以利用纯合成数据来教这些模型克服这些缺点,而不损害它们的零射击能力。我们贡献了合成视觉概念(SyViC)——一个百万规模的合成数据集和数据生成代码库,允许生成额外的合适数据,以提高VLC的理解和VL模型的组合推理。此外,我们提出了一个通用的VL调优策略,以有效地利用SyViC来实现这些改进。我们在VL- checklist、Winoground和ARO基准上进行了广泛的实验和实验,结果表明,使用合成数据来适应强预训练的VL模型可以显著提高VLC的理解能力(例如,在ARO上提高9.9%,在VL- checklist上提高4.3%),其零射击精度下降不到1%。
Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot open vocabulary reasoning over (almost arbitrary) natural language prompts. However, recent works have uncovered a fundamental weakness of these models. For example, their difficulty to understand Visual Language Concepts (VLC) that go ‘beyond nouns’ such as the meaning of non-object words (e.g., attributes, actions, relations, states, etc.), or difficulty in performing compositional reasoning such as understanding the significance of the order of the words in a sentence. In this work, we investigate to which extent purely synthetic data could be leveraged to teach these models to overcome such shortcomings without compromising their zero-shot capabilities. We contribute Synthetic Visual Concepts (SyViC) - a million-scale synthetic dataset and data generation codebase allowing to generate additional suitable data to improve VLC understanding and compositional reasoning of VL models. Additionally, we propose a general VL finetuning strategy for effectively leveraging SyViC towards achieving these improvements. Our extensive experiments and ablations on VL-Checklist, Winoground, and ARO benchmarks demonstrate that it is possible to adapt strong pre-trained VL models with synthetic data significantly enhancing their VLC understanding (e.g. by 9.9% on ARO and 4.3% on VL-Checklist) with under 1% drop in their zero-shot accuracy.