Tolerating memory latency through software-controlled pre-execution in simultaneous multithreading processors

Tolerating memory latency through software-controlled pre-execution in simultaneous multithreading processors
复制标题

DOI:
10.1109/isca.2001.937430
复制
发表时间:
2001-06
期刊:
Proceedings 28th Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
C. Luk
C. Luk
中科院分区:
其他
文献类型:
--
作者:
C. Luk

文献摘要

被引文献

相似文献

在许多不规则的应用程序中,难以预测的数据地址使预取无效。在许多情况下,预测这些地址的唯一准确方法是直接执行生成它们的代码。随着多线程体系结构越来越流行,一种有吸引力的方法是在这些机器上使用空闲线程来执行预执行-本质上是推测地址生成和预取的组合行为,以加速主线程。在本文中,我们提出了这样一个预执行技术的同时多线程(SMT)处理器。通过使用软件来控制预执行,我们能够处理一些通常难以预取的最重要的访问模式。与现有的预执行工作相比,我们的技术实现起来要简单得多(例如,不集成预执行结果、不需要缩短用于预执行的程序、以及不需要专门的硬件来在线程产生时复制寄存器值)。因此,只有最小的扩展SMT机器需要支持我们的技术。尽管它的简单性,我们的技术提供了一组不规则的应用程序,这是一个19%的速度比最先进的软件控制的预取的平均加速24%。
Hardly predictable data addresses in many irregular applications have rendered prefetching ineffective. In many cases, the only accurate way to predict these addresses is to directly execute the code that generates them. As multithreaded architectures become increasingly popular, one attractive approach is to use idle threads on these machines to perform pre-execution-essentially a combined act of speculative address generation and prefetching to accelerate the main thread. In this paper we propose such a pre-execution technique for simultaneous multithreading (SMT) processors. By using software to control pre-execution, we are able to handle some of the most important access patterns that are typically difficult to prefetch. Compared with existing work on pre-execution, our technique is significantly simpler to implement (e.g., no integration of pre-execution results, no need of shortening programs for pre-execution, and no need of special hardware to copy register values upon thread spawns). Consequently, only minimal extensions to SMT machines are required to support our technique. Despite its simplicity, our technique offers an average speedup of 24% in a set of irregular applications, which is a 19% speedup over state-of-the-art software-controlled prefetching.