课题基金 / 基金详情

基于统一数学化描述的深度学习系统对抗性现象研究

批准号:
62076213
项目类别:
面上项目
资助金额:
59.0 万元
负责人:
吴保元
学科分类:
机器感知与机器视觉
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
吴保元

项目摘要

结项摘要

相似基金

相关文献

中文摘要
对抗性现象是深度学习系统面临的严重安全威胁之一,即在正常样本上加入人眼难以察觉的恶意噪声,就可明显改变系统的输出。尽管该领域已出现不少对抗攻击和防御方法,但对于对抗性现象的原理依然缺乏深刻认识,因而导致对抗防御方法往往缺乏严格的理论指导,难以真正保证系统的安全。本项目旨在以针对对抗性原理的全面分析为基础,探索构建安全可靠的深度学习系统的有效解决方案。本项目首先提出对抗性现象的统一数学化描述,清晰刻画了深度学习的基本要素,即模型、数据和损失函数,与对抗性现象的关系。该描述为训练和测试阶段的对抗攻击与防御等研究分支提供了统一视角,还可揭示各分支的内在联系。在该描述的指导下,本项目拟从模型结构、数据特征和对抗训练机制三个维度分析对抗性原理,并将探索基于神经网络架构搜索的鲁棒模型结构搜索方法、基于强化学习的鲁棒特征分离策略、和基于对抗样本分布规律对抗棒训练方法,以此构建全面的对抗防御体系。
英文摘要
Adversarial phenomenon is one of the most serious security threats for deep learning systems. It describes that adding an imperceptible malicious noise onto one benign example could easily fool deep learning models. Although many methods of adversarial attacks and defenses have been developed, there is still lack of a deep understanding of the intrinsic reason of adversarial phenomenon. Consequently, current adversarial defenses are often without solid theoretical support, such that the security of deep learning systems cannot be truly guaranteed. The goal of this project is to explore effective methods of constructing secure and trusty deep learning systems, based on a comprehensive analysis of the intrinsic reason of adversarial phenomenon. To this end, we firstly propose a unified mathematical description about adversarial phenomenon, which clearly describes the relationship between adversarial phenomenon and three basic concepts of deep learning systems, including models, data and loss functions. It not only provides a unified perspective for different branches, including adversarial attacks and defenses in both the training and testing stage, but also builds intrinsic connections among these branches. According to this unified description, the project plans to analyze the intrinsic reason of adversarial phenomenon from three perspectives of model architectures, data and adversarial training mechanisms. We also plan to explore the learning method of robust model architectures based on neural architecture search, the separation strategy of robust features based on reinforcement learning, as well as the robust training method utilizing the distribution of adversarial examples, in order to construct comprehensive adversarial defense systems.
本项目以构建安全可靠的深度学习系统为研究目标,围绕深度学习系统的对抗性现象开展系统性研究。项目组在研究周期内严格遵循原定计划,取得了多项重要进展:在对抗攻击方面,提出了基于样本依赖的隐蔽后门攻击方法、基于二进制整数规划的权重攻击算法、针对目标检测的平行矩形翻转黑盒攻击方法,以及结合元学习的物理对抗攻击方案;在对抗防御方面,从预训练、训练、训练后和推理四个阶段构建了全面的防御体系,提出了基于特征差异的后门样本过滤方法、改进的对抗训练策略、高效的后门模型修复算法(FT-SAM、SAU、NPD)以及随机噪声防御(RND)等创新性解决方案;此外,项目还构建了当前最全面的后门学习基准,包含20种攻击方法和32种防御方法,提供了11000对攻防测试。项目组在研究期间共发表高水平论文37篇,其中包括TPAMI 2篇、ICCV 7篇、NeurIPS 10篇、CVPR 6篇等,相关研究成果为提升深度学习系统的安全性和可靠性提供了重要支持。
国内基金
海外基金