联合视觉语义和文字信息情感挖掘的社交媒体图像理解与评论方法研究
批准号:
62102384
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
方山城
依托单位:
学科分类:
计算机图像视频处理与多媒体技术
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
方山城
中文摘要
社交媒体图像理解与评论旨在以类人方式为图像产生用户评论,是通过引导用户舆论实现有效舆论治理的重要环节。现有方法主要针对图像中的客观事实生成描述,无法产生带有主观情感的高质量评论。为此,本项目拟联合图像中的视觉语义和文字信息充分挖掘图像情感属性,进而生成真实可控的评论。首先,针对难以统一提取图像中异质性的视觉语义和文字信息问题,研究统一框架下异质性信息联合增强的方法,实现精准高效的内容检测与识别;其次,针对难以全面刻画图像中客观事实、情感观点以及背景知识的问题,研究从文字信息中挖掘情感观点、并融合多源内容构建观点场景图的信息表示方法;最后,针对难以真实可控地生成具有情感色彩评论的问题,研究情感条件受控下自编码语言模型结合生成对抗学习产生可控评论的方法。本项目将在社交媒体真实应用场景下验证方案的性能,以期推动图像评论在基础理论及关键技术方面的进展,为舆论引导与治理提供理论支撑及技术保障。
英文摘要
Image understanding and comment on social media aim to generate user comments in a human-like way, which is an important procedure to realize effective public opinion governance through guiding public opinion. Current methods typically generate descriptions for objective facts in images, which is difficult to obtain high-quality comments with subjective sentiment. To this end, this project aims to fully mine the image sentiment jointly with the visual semantics and text information for realistic and controllable comment generation from the following aspects: firstly, for the problem that the heterogeneous visual semantics and text information in the image are difficult to be extracted uniformly, a unified framework would be studied which jointly enhances the heterogeneous information, realizing accurate and efficient detection and recognition; Secondly, for the problem that it is hard to comprehensively represent the objective fact, sentiment opinion and background knowledge in images, we would propose to mine the sentiment opinion from text content, and then integrate multi-source content to construct opinion scene graph for information representation. Finally, for the problem that it is challenging to generate subjective and emotional comments realistically and naturally, a method of non-autoregressive language model under the control of sentiment conditions would be studied, which applies generative adversarial learning to generate realistic and controllable comments. This project will verify the performance of the proposed solution in real application scenarios of social media. It could promote the progress of image comment in basic theories and key techniques, and provide theoretical and technical support for public opinion guidance and governance.
本项目旨在充分利用图像中的视觉语义和文字信息来挖掘图像情感属性,实现深度理解并生成高质量评论,为抹黑性和误导性的内容舆论引导提供支撑。首先,针对图像信息提取难以统一的问题,从视觉特征和语言建模角度出发,构建了视觉物体与文字字符共享特征的统一提取方法。其次,针对图像中主观情感、客观事实难以获取与描述问题,提出了自适应增强自注意力网络的方法,通过预训练模型分析语义与几何关系,从而丰富视觉关系特征。最后,针对目前难以生成主观情感色彩评论的问题,设计情感导向的变分自编码网络模型,提出基于多模态提示下微调策略的评论生成模型。本项目在国际著名期刊和会议上发表24篇高水平研究论文、授权专利2项、申请专利1项,获得软件著作权1项;参与培养5名博士研究生、2名硕士研究生。本项目成果提出的ABINet系列算法也受到学术界和工业界的一致好评,图像和视频的多样化描述方法在实际业务中得到应用。
国内基金
海外基金