Modularized Textual Grounding for Counterfactual Resilience

Modularized Textual Grounding for Counterfactual Resilience
复制标题

DOI:
10.1109/cvpr.2019.00654
复制
发表时间:
2019-04
期刊:
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Zhiyuan Fang;Shu Kong;Charless C. Fowlkes;Yezhou Yang
Zhiyuan Fang;Shu Kong;Charless C. Fowlkes;Yezhou Yang
中科院分区:
其他
文献类型:
--
作者:
Zhiyuan Fang;Shu Kong;Charless C. Fowlkes;Yezhou Yang

文献摘要

相似文献

计算机视觉应用程序通常需要具有精确,解释性和对反事实输入/查询的弹性的文本接地模块。为了达到较高的基础精确度,当前的文本接地方法在很大程度上依赖大规模的培训数据,并在像素级别上手动注释。此类注释获得昂贵,因此严重缩小了模型的现实应用程序范围。此外,这些方法中的大多数牺牲了可解释性,可推广性,并且忽略了对反事实投入的弹性的重要性。为了解决这些问题,我们提出了一个视觉接地系统,该系统是1)端到端,以弱监督的方式训练,仅具有图像级注释,而2)由于模块化设计,我们可以反作用地反弹。具体而言,我们将文本描述分解为三个层次:实体,语义属性,颜色信息和逐步执行构图接地。我们通过一系列实验验证了我们的模型,并证明了其对最新方法的改进。特别是,我们的模型的性能不仅超过了其他弱/不监督的方法,甚至可以使用强有力的监督方法,而且还可以解释为决策,并且在反事实类中的表现要比所有其他方法都更好。
Computer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding precision, current textual grounding methods heavily rely on large-scale training data with manual annotations at the pixel level. Such annotations are expensive to obtain and thus severely narrow the model's scope of real-world applications. Moreover, most of these methods sacrifice interpretability, generalizability, and they neglect the importance of being resilient to counterfactual inputs. To address these issues, we propose a visual grounding system which is 1) end-to-end trainable in a weakly supervised fashion with only image-level annotations, and 2) counterfactually resilient owing to the modular design. Specifically, we decompose textual descriptions into three levels: entity, semantic attribute, color information, and perform compositional grounding progressively. We validate our model through a series of experiments and demonstrate its improvement over the state-of-the-art methods. In particular, our model's performance not only surpasses other weakly/un-supervised methods and even approaches the strongly supervised ones, but also is interpretable for decision making and performs much better in face of counterfactual classes than all the others.