Towards Captioning an Image Collection from a Combined Scene Graph Representation Approach

Towards Captioning an Image Collection from a Combined Scene Graph Representation Approach
复制标题

通过组合场景图表示方法为图像集合添加字幕

DOI:
10.1007/978-3-031-27077-2_14
复制
发表时间:
2023
期刊:
Lecture Notes in Computer Science book series
影响因子:
--
通讯作者:
Ide Ichiro
Ide Ichiro
中科院分区:
--
文献类型:
--
作者:
Phueaksri Itthisak;Kastner Marc A.;Kawanishi Yasutomo;Komamizu Takahiro;Ide Ichiro

文献摘要

相似文献

来自自然语言处理领域的大多数内容摘要模型总结文档或段落集合的文本内容。相比之下,对图像集合的视觉内容进行总结还没有研究到这种程度。在本文中,我们提出了一个框架,总结了图像集合的视觉内容。其关键思想是收集图像集合中所有图像的场景图,创建组合表示,然后使用场景图字幕模型生成可视化摘要字幕。请注意,这旨在将所有图像的共同内容总结在单个标题中,而不是单独描述每个图像。在将图像集合的所有场景图聚合成单个场景图之后,我们通过使用额外的概念泛化组件对其进行规范化。该组件基于词嵌入技术,使用ConceptNet选择每个子图中的公共概念。最后,我们通过从概念泛化组件中用一个共同的概念替换一个特定的名词短语来改进字幕结果。我们使用图像分类和图像标题检索技术,基于MS-COCO数据集构建了一个用于此任务的数据集。该数据集上的所提出的方法的评估显示出良好的性能。
Most content summarization models from the field of natural language processing summarize the textual contents of a collection of documents or paragraphs. In contrast, summarizing the visual contents of a collection of images has not been researched to this extent. In this paper, we present a framework for summarizing the visual contents of an image collection. The key idea is to collect the scene graphs for all images in the image collection, create a combined representation, and then generate a visually summarizing caption using a scene-graph captioning model. Note that this aims to summarize common contents across all images in a single caption rather than describing each image individually. After aggregating all the scene graphs of an image collection into a single scene graph, we normalize it by using an additional concept generalization component. This component selects the common concept in each sub-graph with ConceptNet based on word embedding techniques. Lastly, we refine the captioning results by replacing a specific noun phrase with a common concept from the concept generalization component to improve the captioning results. We construct a dataset for this task based on the MS-COCO dataset using techniques from image classification and image-caption retrieval. An evaluation of the proposed method on this dataset shows promising performance.