基于场景文字的图像描述生成关键技术研究
批准号:
62302255
项目类别:
青年科学基金项目(C类)
资助金额:
10.0 万元
负责人:
王晶
依托单位:
学科分类:
计算机图像视频处理与多媒体技术
结题年份:
2024
批准年份:
2023
项目状态:
已结题
项目参与者:
王晶
中文摘要
图像描述生成任务旨在生成与图像内容相符且逻辑合理的文本描述。其横跨视觉和语言两大模态,可应用于机器人聊天、视障人群辅助等场景,具有重要的研究价值和广阔的应用前景。然而,现有算法大多不具备读取图像中场景文字的能力,导致生成的描述文本容易遗漏图像中的重要信息。为此,本项目拟以含有场景文字的自然场景图像为研究对象,融合计算机视觉、自然语言处理等领域知识,借助预训练大模型、自监督学习、注意力机制等最新技术手段,系统地开展基于场景文字的图像描述生成研究。具体地,拟提出基于视觉关系和语义关系的多模态表征学习方法、文本描述中场景文字的定位与表征选取方法、基于对比学习的场景文字表征优化方法和自回归与非自回归融合的语言模型等,以提高多模态表示的表征能力及描述生成模型的鲁棒性。本项目有望突破现有图像描述生成技术的瓶颈,为多模态内容的知识分析、智能生成等研究和应用提供理论支持和技术支撑。
英文摘要
Image captioning aims to generate textual descriptions that accurately reflect the content of an image in a logical and coherent manner. It bridges the two modalities of vision and language and has a wide range of applications, such as in robot chatting and providing visual aids for people with visual impairments. Therefore, it has important research value and promising application prospects. However, most existing approaches do not have the ability to read scene text in images, resulting in the lack of important information from the image in the description. To address this, this project aims to integrate knowledge from fields such as computer vision and natural language processing, and systematically conduct research on text-based image captioning using techniques such as pretrained large models, self-supervised learning, and attention mechanisms. Specifically, we plan to propose a multi-modal representation learning method based on visual and semantic relationships, a method for locating and selecting representations of scene text in text description, a method for optimizing scene text representation based on contrastive learning, and a language model that combines autoregressive and non-autoregressive methods to improve the representation ability of multi-modal inputs and the robustness of captioning models. This project is expected to break through the bottleneck of current image captioning technology and provide both theoretical and technical support for research and applications in knowledge analysis and artificial intelligent generation of multimodal content.
本项目围绕基于场景文字的图像描述生成任务展开研究,旨在解决现有算法在生成描述文本时容易遗漏图像中场景文字信息的问题。通过融合计算机视觉与自然语言处理领域的最新进展,课题组成功完成了预期研究内容,并取得了一系列创新性成果。具体包括:1)提出了基于视觉关系和语义关系的多模态表征学习方法,显著提升了多模态数据的表征能力;2)设计了文本描述中场景文字的定位与表征选取方法,增强了场景文字表征的鲁棒性;3)提出了基于自批判训练方法的图像描述生成模型,显著提高了生成描述的质量;4)针对遥感图像含多尺度目标的问题和可训练图文对数量不足的问题,提出了基于 CLIP 潜空间和多尺度分组Transformer的图像描述生成算法。依托本项目,课题组已在国际期刊IEEE TGRS(中科院一区)发表学术论文1篇。课题组成员积极参加国内及国际学术会议,并受邀担任相关领域国际顶级会议审稿人。本项目支持博士后1名,协助培养博士研究生1名。本项目的完成为多模态内容的知识分析与智能生成提供了重要的理论支持和技术支撑,推动了图像描述生成技术的实用化进程。
国内基金
海外基金