Semantically Distributed Robust Optimization for Vision-and-Language Inference

Semantically Distributed Robust Optimization for Vision-and-Language Inference
复制标题

DOI:
10.18653/v1/2022.findings-acl.118
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Tejas Gokhale;A. Chaudhary;Pratyay Banerjee;Chitta Baral;Yezhou Yang
Tejas Gokhale;A. Chaudhary;Pratyay Banerjee;Chitta Baral;Yezhou Yang
中科院分区:
其他
文献类型:
--
作者:
Tejas Gokhale;A. Chaudhary;Pratyay Banerjee;Chitta Baral;Yezhou Yang

文献摘要

相似文献

对视觉和语言模型的分析揭示了它们在语言现象下的脆弱性,例如释义,否定,文本范围和具有同义词或反义词的单词替代。而数据增强技术旨在减轻与这些失败模式相对的方法,这些方法可以将这些知识整合到训练管道中。在分布式强大的优化设置中设置语言转换,以及在推理过程中利用这些转换的结合技术。具有图像(NLVR^2)和视频(小提琴)在基准数据集上的典范表明了性能的改进,以及对对抗性攻击的稳健性,对对对抗性的攻击。对Binary VQA进行了binary VQA探索此类概述的综合性,该概述了该方法的其他任务。
Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with synonyms or antonyms.While data augmentation techniques have been designed to mitigate against these failure modes, methods that can integrate this knowledge into the training pipeline remain under-explored.In this paper, we present SDRO, a model-agnostic method that utilizes a set linguistic transformations in a distributed robust optimization setting, along with an ensembling technique to leverage these transformations during inference.Experiments on benchmark datasets with images (NLVR^2) and video (VIOLIN) demonstrate performance improvements as well as robustness to adversarial attacks.Experiments on binary VQA explore the generalizability of this method to other V&L tasks.