Sim VQA: Exploring Simulated Environments for Visual Question Answering

Sim VQA: Exploring Simulated Environments for Visual Question Answering
复制标题

DOI:
10.1109/cvpr52688.2022.00500
复制
发表时间:
2022-03
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Paola Cascante-Bonilla;Hui Wu;Letao Wang;R. Feris;Vicente Ordonez
Paola Cascante-Bonilla;Hui Wu;Letao Wang;R. Feris;Vicente Ordonez
中科院分区:
其他
文献类型:
--
作者:
Paola Cascante-Bonilla;Hui Wu;Letao Wang;R. Feris;Vicente Ordonez

文献摘要

被引文献

相似文献

VQA的现有工作探索数据增强,通过干扰数据集中的图像或修改现有的问题和答案来实现更好的泛化。虽然这些方法表现出良好的性能,但问题和答案的多样性受到可用图像的限制。在这项工作中,我们探索使用合成计算机生成的数据来完全控制视觉和语言空间,使我们能够提供更多样化的场景。我们量化了利用真实VQA合成数据的有效性。通过利用3D和物理模拟平台,我们提供了一个管道来生成合成数据,以扩展和替换特定类型的问题和答案,而不会暴露真实图像中可能存在的敏感或个人数据。我们提供了一个全面的分析,同时扩展现有的超现实数据集用于VQA。我们还提出了特征交换(F-SWAP)——我们在训练期间随机切换对象级特征,使VQA模型更具域不变性。我们表明,F-SWAP可以有效地改进真实图像上的VQA模型,而不会影响其回答数据集中现有问题的准确性。
Existing work on VQA explores data augmentation to achieve better generalization by perturbing images in the dataset or modifying existing questions and answers. While these methods exhibit good performance, the diversity of the questions and answers are constrained by the available images. In this work we explore using synthetic computer-generated data to fully control the visual and language space, allowing us to provide more diverse scenarios. We quantify the effectiveness of leveraging synthetic data for real-world VQA. By exploiting 3D and physics simulation platforms, we provide a pipeline to generate synthetic data to expand and replace type-specific questions and answers without risking exposure of sensitive or personal data that might be present in real images. We offer a comprehensive analysis while expanding existing hyper-realistic datasets to be usedfor VQA. We also propose Feature Swapping (F-SWAP) - where we randomly switch object-level features during training to make a VQA model more domain invariant. We show that F-SWAP is effective for improving VQA models on real images without compromising on their accuracy to answer existing questions in the dataset.