Architecture Support for Improving Bulk Memory Copying and Initialization Performance

Architecture Support for Improving Bulk Memory Copying and Initialization Performance
复制标题

用于提高批量内存复制和初始化性能的架构支持

DOI:
10.1109/pact.2009.31
复制
发表时间:
2009
期刊:
2009 18th International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
R. Iyer
R. Iyer
中科院分区:
--
文献类型:
--
作者:
Xiaowei Jiang;Yan Solihin;Li Zhao;R. Iyer

文献摘要

被引文献

相似文献

大量内存复制和初始化是当前计算机系统中用户应用程序和操作系统执行的最普遍的操作之一。虽然当前许多系统依赖于加载和存储的循环,但有人建议引入单个指令来执行批量内存复制。虽然这样的指令可以提高性能,因为生成更少的TLB和缓存访问,并且需要更少的管道资源,在本文中,我们表明,显著提高性能的关键是消除遵循指令的代码的管道和缓存瓶颈。我们表明瓶颈的产生是由于(1)被复制指令阻塞的管道,(2)由于在等待复制完成时依赖指令停滞而导致的关键路径延长,以及(3)无法(分别)指定源和目标区域的可缓存性。我们提出FastBCI,一种架构支持,实现批量复制/初始化指令的粒度效率,但没有管道和缓存瓶颈。当应用于操作系统内核缓冲区管理时,我们发现FastBCI平均达到23%到32%的加速比,这大约是替代方案的3 -4倍,是高度乐观的DMA的1.5 -2x,并且没有设置和中断开销。
Bulk memory copying and initialization is one of the most ubiquitous operations performed in current computer systems by both user applications and Operating Systems. While many current systems rely on a loop of loads and stores, there are proposals to introduce a single instruction to perform bulk memory copying. While such an instruction can improve performance due to generating fewer TLB and cache accesses, and requiring fewer pipeline resources, in this paper we show that the key to significantly improving the performance is removing pipeline and cache bottlenecks of the code that follows the instructions. We show that the bottlenecks arise due to (1) the pipeline clogged by the copying instruction, (2) lengthened critical path due to dependent instructions stalling while waiting for the copying to complete, and (3) the inability to specify (separately) the cacheability of the source and destination regions. We propose FastBCI, an architecture support that achieves the granularity efficiency of a bulk copying/ initialization instruction, but without its pipeline and cache bottlenecks. When applied to OS kernel buffer management, we show that on average FastBCI achieves anywhere between 23% to 32% speedup ratios, which is roughly 3x-4x of an alternative scheme, and 1.5x-2x of a highly optimistic DMA with zero setup and interrupt overheads.