Depth-2 neural networks under a data-poisoning attack

Depth-2 neural networks under a data-poisoning attack
复制标题

DOI:
10.1016/j.neucom.2023.02.034
复制
发表时间:
2020-05
期刊:
影响因子:
6
通讯作者:
Anirbit Mukherjee-;Ramchandran Muthukumar
Anirbit Mukherjee-;Ramchandran Muthukumar
中科院分区:
计算机科学2区
文献类型:
--
作者:
Anirbit Mukherjee-;Ramchandran Muthukumar

文献摘要

相似文献

在这项工作中,我们研究了在回归设置中训练浅层神经网络时防御数据中毒攻击的可能性。我们专注于为一类深度为2的有限宽度神经网络(包括单过滤器卷积网络)使用可实现的标签进行监督学习。在这类网络中,我们试图在训练过程中,在存在对真实标签进行随机、有界和加性对抗性扭曲的恶意预言机的情况下,学习生成标签的真实网络权重。对于我们构建的无梯度随机算法,我们证明了对抗性攻击的幅度,权重近似精度和所提出的算法实现的置信度之间的最坏情况下的接近最优的权衡。由于我们的算法使用mini-batch,我们分析了mini-batch大小如何影响收敛。我们还展示了如何利用外层权重的缩放来对抗对真实标签的数据中毒攻击,这取决于攻击概率。最后,我们给出了实验证据,证明我们的算法如何在不同的输入数据分布下优于随机梯度下降,包括重尾分布的实例。
In this work, we study the possibility of defending against data-poisoning attacks while training a shallow neural network in a regression setup. We focus on doing supervised learning with realizable labels for a class of depth-2 finite-width neural networks, which includes single-filter convolutional networks. In this class of networks, we attempt to learn the true network weights generating the labels in the presence of a malicious oracle doing stochastic, bounded and additive adversarial distortions on the true labels, during training. For the gradient-free stochastic algorithm that we construct, we prove worst-case near-optimal trade-offs among the magnitude of the adversarial attack, the weight approximation accuracy, and the confidence achieved by the proposed algorithm. As our algorithm uses mini-batching, we analyze how the mini-batch size affects convergence. We also show how to utilize the scaling of the outer layer weights to counter data-poisoning attacks on true labels depending on the probability of attack. Lastly, we give experimental evidence demonstrating how our algorithm outperforms stochastic gradient descent under different input data distributions, including instances of heavy-tailed distributions.