Why nothing matters: the impact of zeroing

Why nothing matters: the impact of zeroing
复制标题

为什么一切都不重要:归零的影响

DOI:
10.1145/2048066.2048092
复制
发表时间:
2011
期刊:
ACM Trans. Program. Lang. Syst.
影响因子:
--
通讯作者:
K. McKinley
K. McKinley
中科院分区:
--
文献类型:
--
作者:
Xi Yang;S. Blackburn;Daniel Frampton;Jennifer B. Sartor;K. McKinley

文献摘要

被引文献

相似文献

内存安全可防止可能危及程序正确性和安全性的内存意外和恶意误用。内存安全的一个关键因素是零初始化。零初始化的直接成本高得令人惊讶:高达12.7%,在IA32架构上的高性能虚拟机上,平均成本从2.7%到4.5%不等。由于其存储器带宽需求和高速缓存移位效应,零初始化还会产生间接成本。现有的虚拟机或者:a)通过在大块中清零来最小化直接成本,或者b)通过在分配序列中清零来最小化间接成本,从而减少高速缓存置换和带宽。本文对两种广泛使用的零初始化设计进行了评估,表明它们在实现非常相似的性能方面做出了不同的权衡。我们的分析启发了三个更好的设计:(1)使用高速缓存旁路(非临时)指令的批量清零,以同时降低直接和间接清零的成本;(2)利用并行硬件将工作移出应用程序的关键路径的并发非临时批量清零;以及(3)自适应清零,它根据可用的硬件并行度在(1)和(2)之间动态选择。新的软件策略提供的加速有时比直接开销更大,使总体性能平均提高3%。我们的发现需要额外的优化和微体系结构支持。
Memory safety defends against inadvertent and malicious misuse of memory that may compromise program correctness and security. A critical element of memory safety is zero initialization. The direct cost of zero initialization is surprisingly high: up to 12.7%, with average costs ranging from 2.7 to 4.5% on a high performance virtual machine on IA32 architectures. Zero initialization also incurs indirect costs due to its memory bandwidth demands and cache displacement effects. Existing virtual machines either: a) minimize direct costs by zeroing in large blocks, or b) minimize indirect costs by zeroing in the allocation sequence, which reduces cache displacement and bandwidth. This paper evaluates the two widely used zero initialization designs, showing that they make different tradeoffs to achieve very similar performance. Our analysis inspires three better designs: (1) bulk zeroing with cache-bypassing (non-temporal) instructions to reduce the direct and indirect zeroing costs simultaneously, (2) concurrent non-temporal bulk zeroing that exploits parallel hardware to move work off the application's critical path, and (3) adaptive zeroing, which dynamically chooses between (1) and (2) based on available hardware parallelism. The new software strategies offer speedups sometimes greater than the direct overhead, improving total performance by 3% on average. Our findings invite additional optimizations and microarchitectural support.