VisualHow: Multimodal Problem Solving

VisualHow: Multimodal Problem Solving
复制标题

DOI:
10.1109/cvpr52688.2022.01518
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Jinhui Yang;Xianyu Chen;Ming Jiang;Shi Chen;Louis Wang;Qi Zhao
Jinhui Yang;Xianyu Chen;Ming Jiang;Shi Chen;Louis Wang;Qi Zhao
中科院分区:
其他
文献类型:
--
作者:
Jinhui Yang;Xianyu Chen;Ming Jiang;Shi Chen;Louis Wang;Qi Zhao

文献摘要

相似文献

计算机视觉(CV)和自然语言处理(NLP)跨学科研究的最新进展使智能系统能够描述他们所看到的并相应地回答问题。然而,尽管在执行这些视觉语言任务方面显示出有用性,现有的方法仍然难以理解现实生活中的问题(即,如何做某事),并提供逐步解决这些问题的指导。我们的总体目标是开发智能系统来帮助人类进行各种日常活动,我们提出了VisualHow,这是一项自由形式和开放式的研究,专注于理解现实生活中的问题,并通过整合跨多种模式的关键组件来推导其解决方案。我们开发了一个新的数据集,其中包含20,028个现实问题和102,933个步骤,这些步骤构成了它们的解决方案,其中每个步骤都包括一个视觉插图和一个指导问题解决的文本描述。为了更好地理解问题和解决方案,我们还提供了多模态注意力的注释,这些注释将重要的组件定位在模态和解决方案图之间,这些图将不同的步骤封装在结构化表示中。这些数据和注释使得一系列新的视觉语言任务能够解决现实生活中的问题。通过对代表性模型的广泛实验,我们证明了它们在训练和测试新任务模型上的有效性,并且通过学习有效的注意机制有很大的改进空间。我们的数据集和模型可在https://github.com/formidify/VisualHow上获得。
Recent progress in the interdisciplinary studies of computer vision (CV) and natural language processing (NLP) has enabled the development of intelligent systems that can describe what they see and answer questions accordingly. However, despite showing usefulness in performing these vision-language tasks, existing methods still struggle in understanding real-life problems (i.e., how to do something) and suggesting step-by-step guidance to solve them. With an overarching goal of developing intelligent systems to assist humans in various daily activities, we propose VisualHow, a free-form and open-ended research that focuses on understanding a real-life problem and deriving its solution by incorporating key components across multiple modalities. We develop a new dataset with 20,028 real-life problems and 102,933 steps that constitute their solutions, where each step consists of both a visual illustration and a textual description that guide the problem solving. To establish better understanding of problems and solutions, we also provide annotations of multimodal attention that localizes important components across modalities and solution graphs that encapsulate different steps in structured representations. These data and annotations enable a family of new vision-language tasks that solve real-life problems. Through extensive experiments with representative models, we demonstrate their effectiveness on training and testing models for the new tasks, and there is significant scope for improvement by learning effective attention mechanisms. Our dataset and models are available at https://github.com/formidify/VisualHow.