PSSM: achieving secure memory for GPUs with partitioned and sectored security metadata

PSSM: achieving secure memory for GPUs with partitioned and sectored security metadata
复制标题

PSSM:通过分区和扇区安全元数据实现 GPU 安全内存

DOI:
10.1145/3447818.3460374
复制
发表时间:
2021
期刊:
The 35th ACM International Conference on Supercomputing
影响因子:
--
通讯作者:
Zhou, Huiyang
Zhou, Huiyang
中科院分区:
--
文献类型:
--
作者:
Yuan, Shougang;Solihin, Yan;Zhou, Huiyang

文献摘要

参考文献

被引文献

相似文献

本文研究了GPU的安全存储结构,指出传统的CPU安全存储结构不能直接应用于GPU。主要原因包括:(1)访问安全元数据(包括加密计数器、消息认证码(MAC)和完整性树)需要大量的存储器带宽,这可能导致与正常数据访问的严重带宽竞争并降低GPU性能;(2)当代GPU使用分区存储器组织,这导致加密计数器和完整性树的存储和一致性问题,因为不同的分区可能需要更新相同的计数器/完整性树块;以及(3)现有的拆分计数器块组织对分区缓存不友好,分区缓存通常用于GPU中以节省带宽。基于这些观察,我们提出了分区和扇区安全元数据(PSSM),它有两个组件:(a)使用偏移地址(称为本地地址),而不是虚拟或物理地址,生成元数据,以便解决计数器或完整性树存储和一致性问题,以及(B)重新组织安全元数据以使它们对分区高速缓存结构友好,从而减少元数据访问的存储器带宽消耗。通过这些方案,安全GPU内存的性能开销平均从59.22%降低到16.84%。如果只需要内存加密,性能开销从29.53%降低到5.18%。
In this paper, we investigate the secure memory architecture for GPUs and point out that conventional CPU secure memory architecture can not be directly adopted to the GPUs. The key reasons include: (1) accessing the security metadata, including encryption counters, message authentication codes (MACs) and integrity trees, requires significant memory bandwidth, which may lead to severe bandwidth competition with normal data accesses and degrade the GPU performance; (2) contemporary GPUs use partitioned memory organization, which results in storage and coherence problems for encryption counters and integrity trees since different partitions may need to update the same counter/integrity tree blocks; and (3) the existing split-counter block organization is not friendly to sectored caches, which are commonly used in GPU for bandwidth savings. Based on these observations, we propose partitioned and sectored security metadata (PSSM), which has two components: (a) using the offset addresses (referred to as local addresses) within each partition, instead of the virtual or physical addresses, to generate the metadata so as to solve the counter or integrity tree storage and coherence problem and (b) reorganizing the security metadata to make them friendly to the sectored cache structure so as to reduce the memory bandwidth consumption of metadata accesses. With these proposed schemes, the performance overhead of secure GPU memory is reduced from 59.22% to 16.84% on average. If only memory encryption is required, the performance overhead is reduced from 29.53% to 5.18%.
分析 GPU 的安全内存架构
DOI: 10.1109/ispass51385.2021.00017
发表时间: 2021
期刊: IEEE International Symposium on Performance Analysis of Systems and Software
影响因子: --
作者:
Yuan, Shougang;Baskara Yudha, Ardhi Wiratama;Solihin, Yan;Zhou, Huiyang
通讯作者: Zhou, Huiyang
使安全处理器对操作系统和性能友好
DOI: 10.1145/1498690.1498691
发表时间: 2009
期刊: ACM Trans. Archit. Code Optim.
影响因子: --
作者:
Siddhartha Chhabra;Brian Rogers;Yan Solihin;Milos Prvulović
通讯作者: Milos Prvulović
第 55 届年度设计自动化会议论文集
DOI: 10.1145/3195970
发表时间: 1998
期刊: Proceedings of the 55th Annual Design Automation Conference
影响因子: --
作者:
J. Rabaey
通讯作者: J. Rabaey
多处理器计算机系统模型中的交错内存带宽
DOI: 10.1109/tc.1979.1675436
发表时间: 1979
影响因子: 3.7
作者:
B. R. Rau
通讯作者: B. R. Rau
DOI: --
发表时间: 2020
期刊: --
影响因子: --
作者:
T. Hunt;Zhipeng Jia;Vance Miller;Ariel Szekely;Yige Hu;C. Rossbach;Emmett Witchel
通讯作者: T. Hunt;Zhipeng Jia;Vance Miller;Ariel Szekely;Yige Hu;C. Rossbach;Emmett Witchel