Address-value delta (AVD) prediction: increasing the effectiveness of runahead execution by exploiting regular memory allocation patterns

Address-value delta (AVD) prediction: increasing the effectiveness of runahead execution by exploiting regular memory allocation patterns
复制标题

地址值增量 (AVD) 预测:通过利用常规内存分配模式提高超前执行的有效性

DOI:
10.1109/micro.2005.11
复制
发表时间:
2005
期刊:
38th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO'05)
影响因子:
--
通讯作者:
Y. Patt
Y. Patt
中科院分区:
--
文献类型:
--
作者:
O. Mutlu;Hyesoon Kim;Y. Patt

文献摘要

被引文献

相似文献

虽然运行前执行在并行化独立的长延迟缓存缺失方面是有效的,但它无法并行化依赖的长延迟缓存缺失。为了克服这一限制,本文提出了一种新的技术,地址值增量(AVD)预测。AVD预测器跟踪有效地址和数据值之间的算术差(即增量)是稳定的地址(指针)加载指令。如果这样的加载指令在运行前执行期间导致长延迟缓存丢失,则通过从其有效地址中减去稳定增量来预测其数据值。此预测支持相关指令的预执行,包括导致长延迟缓存丢失的加载指令。我们描述了如何,为什么,以及什么样的负载AVD预测工作,并评估了一个可实现的AVD预测器的设计权衡。我们的分析表明,由于数据结构在内存中分配的方式的模式,稳定的avd存在。我们的结果表明,使用一个简单的、16个条目的AVD预测器来增加一个预跑处理器,可以使一组指针密集型应用程序的平均执行时间提高12.1%。
While runahead execution is effective at parallelizing independent long-latency cache misses, it is unable to parallelize dependent long-latency cache misses. To overcome this limitation, this paper proposes a novel technique, address-value delta (AVD) prediction. An AVD predictor keeps track of the address (pointer) load instructions for which the arithmetic difference (i.e., delta) between the effective address and the data value is stable. If such a load instruction incurs a long-latency cache miss during runahead execution, its data value is predicted by subtracting the stable delta from its effective address. This prediction enables the pre-execution of dependent instructions, including load instructions that incur long-latency cache misses. We describe how, why, and for what kind of loads AVD prediction works and evaluate the design tradeoffs in an implementable AVD predictor. Our analysis shows that stable AVDs exist because of patterns in the way data structures are allocated in memory. Our results show that augmenting a runahead processor with a simple, 16-entry AVD predictor improves the average execution time of a set of pointer-intensive applications by 12.1%.