The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book Narratives

The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book Narratives
复制标题

阴沟的惊人奥秘:在漫画书叙述中的面板之间进行推论

DOI:
--
复制
发表时间:
2016
期刊:
Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
L. Davis
L. Davis
中科院分区:
--
文献类型:
--
作者:
Mohit Iyyer;Varun Manjunatha;Anupam Guha;Yogarshi Vyas;Jordan L. Boyd;Hal Daumé;L. Davis

文献摘要

被引文献

相似文献

视觉叙事通常是明确的信息和明智的遗漏的结合,依赖于观众提供缺失的细节。在漫画中,大多数时间和空间的运动都隐藏在面板之间的水槽中。为了了解故事,读者通过一个称为闭合的过程推断看不见的动作,从而将面板逻辑地连接在一起。虽然计算机现在可以描述自然图像的内容,但在本文中,我们研究它们是否可以理解漫画书面板中的风格化艺术品和对话所传达的封闭驱动的叙述。我们收集了一个数据集,COMICS,由超过120万个面板(120 GB)与自动文本框转换配对组成。对漫画的深入分析表明,文本和图像都不能单独讲述漫画故事,因此计算机必须理解这两种模式才能跟上情节。我们介绍了三个完形填空式的任务,要求模型预测的叙事和字符为中心的方面,一个面板给定的n个前面的面板作为背景。各种深度神经架构在这些任务上的表现低于人类基线,这表明COMICS包含视觉和语言的基本挑战。
Visual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the gutters between panels. To follow the story, readers logically connect panels together by inferring unseen actions through a process called closure. While computers can now describe the content of natural images, in this paper we examine whether they can understand the closure-driven narratives conveyed by stylized artwork and dialogue in comic book panels. We collect a dataset, COMICS, that consists of over 1.2 million panels (120 GB) paired with automatic textbox transcriptions. An in-depth analysis of COMICS demonstrates that neither text nor image alone can tell a comic book story, so a computer must understand both modalities to keep up with the plot. We introduce three cloze-style tasks that ask models to predict narrative and character-centric aspects of a panel given n preceding panels as context. Various deep neural architectures underperform human baselines on these tasks, suggesting that COMICS contains fundamental challenges for both vision and language.