融合跨模态深度分析与迭代优化的图像描述自动生成方法研究
批准号:
61976057
项目类别:
面上项目
资助金额:
61.0 万元
负责人:
张玥杰
依托单位:
学科分类:
自然语言处理
结题年份:
2023
批准年份:
2019
项目状态:
已结题
项目参与者:
张玥杰
中文摘要
本项目针对跨模态图像描述生成方法进行深入研究,探索、建立和完善面向大规模多模态信息源并融合跨模态深度分析与迭代优化机制的图像描述自动生成基本理论、方法和实现。通过融合多模态信息源的深度理解、分析与优化关键技术,建立有效适用的图像描述自动生成机制,为图像视觉内容的深层次理解提供高效、精确、易用的获取合理语义信息的智能支持。其中,采取图像多模态内容展现形式作为重要基础信息源,有效利用图像视觉信息与文本描述的多模态深度特征属性表征与融合、跨模态深度主题关联分析与学习、跨模态深度语义相关性学习与融合、融合多模态注意力的跨模态描述生成建模、及融合迭代优化的跨模态深度精炼与优选等多种跨模态深层次分析、处理与优化机制,并针对面向跨模态描述生成所需的深度神经网络与多模态注意力相关理论与工作机制进行深入探索,以期构成功能更加强壮且性能更加完善的图像描述自动生成架构,形成视觉内容描述跨模态自动生成的新形态。
英文摘要
This project focuses on the research for image captioning with cross-modal deep analysis and iterative optimization, so as to explore, establish and improve the basic theory, approach and implementation for image captioning that fuses large-scale multimodal information sources, cross-modal deep analysis and iterative optimization mechanisms. It aims to construct a more effective and appropriate image captioning scheme by integrating the deep understanding, analysis and optimization techniques for multimodal information sources. Thus the intelligent support for efficient, precise and easy-to-use acquirement of reasonable semantic information can be provided for deep-level visual content understanding. Visual information users can promote their efficiency for obtaining beneficial and valuable information from large-scale image data. Multiple new key techniques for image captioning are deeply studied and developed to facilitate more effective caption generation with the stronger function and more perfect performance. They focus on multiple cross-modal deep-level analysis, processing and optimization mechanisms, including multimodal deep feature attribute representation and fusion, cross-modal deep topic association analysis and learning, cross-modal deep semantic correlation learning and fusion, cross-modal captioning modeling with multimodal attention, and cross-modal deep refinement and optimization with iterative learning, etc. Meanwhile, the related theory and working mechanisms for deep neural network and multimodal attention are also deeply explored. All these research work aims to establish a stronger and better cross-modal image captioning architecture, and construct a new form for visual content captioning.
期刊论文列表
专著列表
科研奖励列表
会议论文列表
专利列表
登录
查看更多内容
Deep Cross-Modal Face Naming for People News Retrieval
用于人物新闻检索的深度跨模态人脸命名
DOI:
10.1109/tkde.2019.2948875
发表时间:
2019-10
期刊:
IEEE Transactions on Knowledge and Data Engineering
影响因子:
8.9
作者:
[Yong Tian, Lian Zhou, Yuejie Zhang, Tao Zhang, Weiguo Fan]
通讯作者:
Weiguo Fan
DOI:
--
发表时间:
2022
期刊:
中文信息学报
影响因子:
作者:
[周练, 俞国瑞, 张玥杰, 冯瑞, 张涛, 张晓波]
通讯作者:
张晓波
DOI:
10.1016/j.patcog.2021.108258
发表时间:
2021-08-29
期刊:
PATTERN RECOGNITION
影响因子:
8
作者:
[Miao, Shuyu, Du, Shanshan, Fan, Weiguo]
通讯作者:
Fan, Weiguo
DOI:
10.1016/j.patcog.2021.108005
发表时间:
2021-10
期刊:
Pattern recognition
影响因子:
8
作者:
[Hou J, Xu J, Jiang L, Du S, Feng R, Zhang Y, Shan F, Xue X]
通讯作者:
Xue X
DOI:
10.1109/tcsvt.2021.3058626
发表时间:
2021-02
期刊:
IEEE Transactions on Circuits and Systems for Video Technology
影响因子:
8.4
作者:
[Y. Zheng;Yuejie Zhang;Rui Feng;Tao Zhang;Weiguo Fan]
通讯作者:
Y. Zheng;Yuejie Zhang;Rui Feng;Tao Zhang;Weiguo Fan
共 10 条
基于视觉-语言模型的医学影像序列分析与智能诊断关键技术研究
-
批准号:--
-
项目类别:省市级项目
-
资助金额:0.0万元
-
批准年份:2025
-
负责人:张玥杰
-
依托单位:
融合多模态文本关联分析与挖掘的跨媒体社会图像检索方法研究
-
批准号:61572140
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:张玥杰
-
依托单位:
面向英汉双向跨语言图像检索的文本分析关键技术研究
-
批准号:61170095
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2011
-
负责人:张玥杰
-
依托单位:
面向英汉双向跨语言信息检索的若干自然语言处理底层关键技术研究
-
批准号:60773124
-
项目类别:面上项目
-
资助金额:24.0万元
-
批准年份:2007
-
负责人:张玥杰
-
依托单位:
面向数据的英汉双向跨语言信息检索关键技术研究
-
批准号:60203010
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2002
-
负责人:张玥杰
-
依托单位:
国内基金
海外基金