Semantically Distributed Robust Optimization for Vision-and-Language Inference
Semantically Distributed Robust Optimization for Vision-and-Language Inference
复制标题
DOI:
10.18653/v1/2022.findings-acl.118
复制
发表时间:
2021-10
期刊:
影响因子:
--
通讯作者:
Tejas Gokhale;A. Chaudhary;Pratyay Banerjee;Chitta Baral;Yezhou Yang
中科院分区:
文献类型:
--
作者:
Tejas Gokhale;A. Chaudhary;Pratyay Banerjee;Chitta Baral;Yezhou Yang
Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with synonyms or antonyms.While data augmentation techniques have been designed to mitigate against these failure modes, methods that can integrate this knowledge into the training pipeline remain under-explored.In this paper, we present SDRO, a model-agnostic method that utilizes a set linguistic transformations in a distributed robust optimization setting, along with an ensembling technique to leverage these transformations during inference.Experiments on benchmark datasets with images (NLVR^2) and video (VIOLIN) demonstrate performance improvements as well as robustness to adversarial attacks.Experiments on binary VQA explore the generalizability of this method to other V&L tasks.