课题基金 / 基金详情

基于因果分析的视觉内容描述研究

批准号:
62102070
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
宾燚
依托单位:
学科分类:
计算机图像视频处理与多媒体技术
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
宾燚

项目摘要

结项摘要

相似基金

相关文献

中文摘要
近年来,得益于深度学习强大的特征感知和表示能力,视觉内容描述研究取得了显著进展。然而,面对物理世界中的有偏数据,深度学习方法由于缺乏归纳一般性规则和推理能力,无法正确区分有用的特征模式和无用的偏差,从而导致模型鲁棒性和泛化能力较差。基于此,本项目引入因果分析方法,结合深度学习技术研究视觉内容描述。本项目重点研究“基于图像语义编辑的视觉内容描述方法”、“基于因果混杂因子分析的视觉内容描述方法”和“基于场景图的反事实视觉内容描述方法”三方面内容,分别从浅层数据分布干预、中层潜在因果混杂因子分析以及深层的反事实因果推理三个层次,层层递进,提升视觉内容描述系统性能。本项目的研究成果能为如何建立具有简单认知推理能力的视觉内容描述系统提供新方法和新思路,并推动视觉内容理解及语义分析在复杂真实环境的应用。
英文摘要
Visual captioning has made great progress in recent years, thanks to the powerful capacities in feature perception and representation of deep learning. However, deep learning, poor in reasoning and induction, cannot discriminate positive patterns and negative patterns (biases) effectively, resulting in models with poor robustness and generalization for real-world biased data. Towards this end, this project introduces causal analysis and combines it with deep learning for visual captioning. Specifically, this project focuses on the research of three aspects: (1) visual captioning based on image semantics editing, (2) visual captioning via causal confounder analysis, and (3) scene graph based counterfactual visual captioning. These research aim to improve the effectiveness of visual captioning system from shallow data distribution intervention, middle-level latent causal confounder analysis, and deep causal reasoning with counterfactuals. The research findings of this project can provide new strategies and new ideas for how to build an intelligent visual captioning system with the ability of reasoning, and promote the application of visual content understanding and semantic analysis in complex real-world scenarios.
本项目瞄准视觉内容描述研究中的特征表示学习,旨在通过因果分析方法消除数据中存在的伪关联,探索视觉语义理解本质特征。基于此,本项目开展了视觉语义编辑与合成、因果混杂因子分析和反事实视觉内容描述三方面研究。针对多模态语义编辑合成,提出语义信息不对称、多模态语义可控等代表性方法,实现视觉-语言数据语义可控编辑与生成,从数据支层面为后续因果分析方法提供支撑;在因果混杂因子分析方面,提出多模态因子解耦、跨模态语义一致性建模等方法,消除混在因子在多模态语义表示与理解方面的影响;在反事实视觉内容描述研究方面,本项目从视觉-语言数据假阴性问题出发,提出基于贝叶斯估计的假阴性消除方法,从而构建数据标签反事实假设,使得模型能更好地学习视觉语义本质特征。最后,本项目在艺术绘画视觉内容描述方面开展应用研究,有效探索艺术绘画领域视觉内容本质特征和艺术表现技巧因果关系,实现了视觉语义推理与描述研究。基于这些研究,本项目发表高水平论文11篇,其中CCF A类或中科院一区9篇,第一作者8篇(含共同一作),申请发明专利2项,形成了具有自主知识产权的研究成果,达到预期研究目标。
国内基金
海外基金