Divide and Conquer: Answering Questions with Object Factorization and Compositional Reasoning

Divide and Conquer: Answering Questions with Object Factorization and Compositional Reasoning
复制标题

DOI:
10.48550/arxiv.2303.10482
复制
发表时间:
2023-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Shi Chen;Qi Zhao
Shi Chen;Qi Zhao
中科院分区:
其他
文献类型:
--
作者:
Shi Chen;Qi Zhao

文献摘要

相似文献

人类具有回答不同问题的先天能力,这源于基于语义关系将不同概念关联起来并将困难问题分解为子任务的自然能力。相反,现有的视觉推理方法假设训练样本捕捉每一个可能的对象和推理问题,并依赖于黑盒模型,通常利用统计先验。他们还没有发展出在现实世界中解决新事物或虚假偏见的能力,也没有解释他们决定背后的理由。受人类对视觉世界推理的启发,我们从组合的角度解决了上述挑战,并提出了一个由原则性对象分解方法和新型神经模块网络组成的整体框架。我们的因式分解方法根据对象的关键特征进行分解,并自动导出代表各种对象的原型。通过这些原型编码重要的语义,所提出的网络然后通过在公共语义空间上测量对象的相似性来关联对象,并通过组合推理过程做出决策。它能够回答不同对象的问题,而不管它们在训练过程中是否可用,并克服有偏见的问答分布问题。除了增强的泛化能力,我们的框架还提供了一个可解释的接口,用于理解模型的决策过程。我们的代码可在https://github.com/szzexpoi/POEM上获得。
Humans have the innate capability to answer diverse questions, which is rooted in the natural ability to correlate different concepts based on their semantic relationships and decompose difficult problems into sub-tasks. On the contrary, existing visual reasoning methods assume training samples that capture every possible object and reasoning problem, and rely on black-boxed models that commonly exploit statistical priors. They have yet to develop the capability to address novel objects or spurious biases in real-world scenarios, and also fall short of interpreting the rationales behind their decisions. Inspired by humans' reasoning of the visual world, we tackle the aforementioned challenges from a compositional perspective, and propose an integral framework consisting of a principled object factorization method and a novel neural module network. Our factorization method decomposes objects based on their key characteristics, and automatically derives prototypes that represent a wide range of objects. With these prototypes encoding important semantics, the proposed network then correlates objects by measuring their similarity on a common semantic space and makes decisions with a compositional reasoning process. It is capable of answering questions with diverse objects regardless of their availability during training, and overcoming the issues of biased question-answer distributions. In addition to the enhanced generalizability, our framework also provides an interpretable interface for understanding the decision-making process of models. Our code is available at https://github.com/szzexpoi/POEM.