Improving Memory Subsystem Performance Using ViVA: Virtual Vector Architecture

Improving Memory Subsystem Performance Using ViVA: Virtual Vector Architecture
复制标题

使用 ViVA 提高内存子系统性能:虚拟向量架构

DOI:
10.1007/978-3-642-00454-4_16
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
K. Yelick
K. Yelick
中科院分区:
--
文献类型:
--
作者:
Joseph Gebis;L. Oliker;J. Shalf;Samuel Williams;K. Yelick

文献摘要

被引文献

相似文献

微处理器时钟频率和内存延迟之间的差异是许多要求苛刻的应用程序运行远低于可实现的峰值性能的主要原因。软件控制的刮刮板存储器,如Cell本地存储器,试图通过精确控制存储器的移动来改善这种差异;然而,scratchpad技术使程序员和编译器面临一个不熟悉和困难的编程模型。在这项工作中,我们提出了虚拟向量架构(ViVA),它将向量计算机的内存语义与软件控制的刮刮板存储器相结合,以提供更有效和实用的延迟隐藏方法。ViVA需要对核心设计进行最小的更改,因此可以很容易地与传统处理器核心集成。为了验证我们的方法,我们在Mambo周期精确的全系统模拟器上实现了ViVA,该模拟器经过仔细校准,以匹配底层PowerPC Apple G5架构上的性能。结果表明,ViVA能够在各种内存访问模式以及两个重要的内存绑定紧凑内核(拐角转换和稀疏矩阵向量乘法)上提供比标量技术显著的性能优势,与标量版本相比实现了2 - 13倍的改进。总的来说,我们初步的ViVA探索指出了一种有前途的方法,可以以最小的设计和复杂性成本,以节能的方式提高领先微处理器上的应用程序性能。
The disparity between microprocessor clock frequencies and memory latency is a primary reason why many demanding applications run well below peak achievable performance. Software controlled scratchpad memories, such as the Cell local store, attempt to ameliorate this discrepancy by enabling precise control over memory movement; however, scratchpad technology confronts the programmer and compiler with an unfamiliar and difficult programming model. In this work, we present the Virtual Vector Architecture (ViVA), which combines the memory semantics of vector computers with a software-controlled scratchpad memory in order to provide a more effective and practical approach to latency hiding. ViVA requires minimal changes to the core design and could thus be easily integrated with conventional processor cores. To validate our approach, we implemented ViVA on the Mambo cycle-accurate full system simulator, which was carefully calibrated to match the performance on our underlying PowerPC Apple G5 architecture. Results show that ViVA is able to deliver significant performance benefits over scalar techniques for a variety of memory access patterns as well as two important memory-bound compact kernels, corner turn and sparse matrix-vector multiplication — achieving 2x–13x improvement compared the scalar version. Overall, our preliminary ViVA exploration points to a promising approach for improving application performance on leading microprocessors with minimal design and complexity costs, in a power efficient manner.