Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU Systems

Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU Systems
复制标题

DOI:
10.1145/3579371.3589080
复制
发表时间:
2022-03
期刊:
Proceedings of the 50th Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
S. B. Dutta;Hoda Naghibijouybari;Arjun Gupta;Nael B. Abu-Ghazaleh;A. Márquez;K. Barker
S. B. Dutta;Hoda Naghibijouybari;Arjun Gupta;Nael B. Abu-Ghazaleh;A. Márquez;K. Barker
中科院分区:
其他
文献类型:
--
作者:
S. B. Dutta;Hoda Naghibijouybari;Arjun Gupta;Nael B. Abu-Ghazaleh;A. Márquez;K. Barker

文献摘要

相似文献

深度学习革命在很大程度上是由gpu和最近的加速器实现的,这使得在可接受的时间内执行计算要求高的训练和推理成为可能。随着机器学习网络的规模和工作负载的不断增加,多gpu机器已经成为高性能计算和云数据中心提供的重要平台。由于这些机器在多个用户之间共享,因此保护应用程序免受潜在攻击变得越来越重要。在本文中,我们探讨了Nvidia的DGX多gpu机器在隐蔽和侧信道攻击中的漏洞。这些机器由许多独立的gpu组成,这些gpu通过定制互连(NVLink)和PCIe连接的组合相互连接。我们对相互连接的缓存层次结构进行逆向工程,并表明一个GPU上的攻击者可能会在另一个GPU的L2缓存上引起争用。我们利用这一观察结果首先开发了跨两个gpu的隐蔽通道攻击,实现了大约4 MB/s的最佳带宽。我们还开发了远程GPU上的初始和探测攻击,允许攻击者恢复另一个工作负载的缓存访问模式。这种访问模式可以用于任何数量的侧信道攻击:我们演示了一种概念验证攻击,该攻击可以高精度地识别在远程GPU上运行的应用程序。我们还开发了一种概念证明攻击来提取机器学习工作负载的超参数。我们的工作首次确定了这些机器对微架构攻击的脆弱性,并可以指导未来的研究以提高其安全性。
The deep learning revolution has been enabled in large part by GPUs, and more recently accelerators, which make it possible to carry out computationally demanding training and inference in acceptable times. As the size of machine learning networks and workloads continues to increase, multi-GPU machines have emerged as an important platform offered on High Performance Computing and cloud data centers. Since these machines are shared among multiple users, it becomes increasingly important to protect applications against potential attacks. In this paper, we explore the vulnerability of Nvidia's DGX multi-GPU machines to covert and side channel attacks. These machines consist of a number of discrete GPUs that are interconnected through a combination of custom interconnect (NVLink) and PCIe connections. We reverse engineer the interconnected cache hierarchy and show that it is possible for an attacker on one GPU to cause contention on the L2 cache of another GPU. We use this observation to first develop a covert channel attack across two GPUs, achieving the best bandwidth of around 4 MB/s. We also develop a prime and probe attack on a remote GPU allowing an attacker to recover the cache access pattern of another workload. This access pattern can be used in any number of side channel attacks: we demonstrate a proof of concept attack that fingerprints the application running on the remote GPU, with high accuracy. We also develop a proof of concept attack to extract hyperparameters of a machine learning workload. Our work establishes for the first time the vulnerability of these machines to microarchitectural attacks and can guide future research to improve their security.