Visualizing and Analyzing the Topology of Neuron Activations in Deep Adversarial Training

Visualizing and Analyzing the Topology of Neuron Activations in Deep Adversarial Training
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Youjia Zhou;Yi Zhou;Jie Ding;Bei Wang
Youjia Zhou;Yi Zhou;Jie Ding;Bei Wang
中科院分区:
其他
文献类型:
--
作者:
Youjia Zhou;Yi Zhou;Jie Ding;Bei Wang

文献摘要

相似文献

众所周知,深度模型容易受到数据对抗性攻击,已经开发了许多对抗性训练技术来提高其对抗性鲁棒性。虽然数据攻击者通过修改数据来攻击模型预测,但他们对模型产生的神经元激活的影响知之甚少,而神经元激活在确定模型的预测和可解释性方面起着至关重要的作用。在这项工作中,我们的目标是发展对抗训练的拓扑理解,以增强其可解释性。我们分析了深度对抗训练产生的数据样本的神经元激活的拓扑结构,特别是映射图。映射图的每个节点表示一个激活簇,如果两个节点对应的簇有非空交集,则它们通过边连接。我们提供了一个交互式的可视化工具,展示了我们的拓扑框架在探索激活空间的效用。我们发现,更强的攻击使数据样本在神经元激活空间中更难以区分,从而导致准确性降低。我们的工具还提供了一种自然的方法来识别脆弱的数据样本,这可能有助于提高模型的鲁棒性。
Deep models are known to be vulnerable to data adversarial attacks, and many adversarial training techniques have been developed to improve their adversarial robustness. While data adversaries attack model predictions through modifying data, little is known about their impact on the neuron activations produced by the model, which play a crucial role in determining the model’s predictions and interpretability. In this work, we aim to de-velop a topological understanding of adversarial training to enhance its interpretability. We analyze the topological structure—in particular, mapper graphs—of neuron activations of data samples produced by deep adversarial training. Each node of a mapper graph represents a cluster of activations, and two nodes are connected by an edge if their corresponding clusters have a nonempty intersection. We provide an interactive visualization tool that demonstrates the utility of our topological framework in exploring the activation space. We found that stronger attacks make the data samples more indistinguishable in the neuron activation space that leads to a lower accuracy. Our tool also provides a natural way to identify the vulnerable data samples that may be useful in improving model robustness.