Bluff: Interactively Deciphering Adversarial Attacks on Deep Neural Networks

Bluff: Interactively Deciphering Adversarial Attacks on Deep Neural Networks
复制标题

DOI:
10.1109/vis47514.2020.00061
复制
发表时间:
2020-09
期刊:
2020 IEEE Visualization Conference (VIS)
影响因子:
--
通讯作者:
Nilaksh Das;Haekyu Park;Zijie J. Wang;Fred Hohman;Robert Firstman;Emily Rogers;Duen Horng Chau
Nilaksh Das;Haekyu Park;Zijie J. Wang;Fred Hohman;Robert Firstman;Emily Rogers;Duen Horng Chau
中科院分区:
其他
文献类型:
--
作者:
Nilaksh Das;Haekyu Park;Zijie J. Wang;Fred Hohman;Robert Firstman;Emily Rogers;Duen Horng Chau

文献摘要

相似文献

深度神经网络(DNN)现在通常用于许多领域。然而,它们容易受到对抗性攻击:精心设计的数据输入扰动,可以欺骗模型做出错误的预测。尽管在开发DNN攻击和防御技术方面进行了大量研究,但人们仍然缺乏对此类攻击如何渗透模型内部的理解。我们提出了海崖,一个交互式系统,用于可视化,表征和破译基于视觉的神经网络的对抗性攻击。海崖允许人们灵活地可视化和比较良性和受攻击图像的激活路径,揭示对抗性攻击对模型造成伤害的机制。海崖是开源的,可以在现代的网络浏览器中运行。
Deep neural networks (DNNs) are now commonly used in many domains. However, they are vulnerable to adversarial attacks: carefully-crafted perturbations on data inputs that can fool a model into making incorrect predictions. Despite significant research on developing DNN attack and defense techniques, people still lack an understanding of how such attacks penetrate a model’s internals. We present Bluff, an interactive system for visualizing, characterizing, and deciphering adversarial attacks on vision-based neural networks. Bluff allows people to flexibly visualize and compare the activation pathways for benign and attacked images, revealing mechanisms that adversarial attacks employ to inflict harm on a model. Bluff is open-sourced and runs in modern web browsers.