Hardware prefetchers for emerging parallel applications

Hardware prefetchers for emerging parallel applications
复制标题

适用于新兴并行应用的硬件预取器

DOI:
--
复制
发表时间:
2012
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
S. Balachandran
S. Balachandran
中科院分区:
--
文献类型:
--
作者:
Biswabandan Panda;S. Balachandran

文献摘要

被引文献

相似文献

硬件预取在过去已经被研究用于多编程工作负载以及GPU。高效的硬件预取器(如基于流或基于GHB的预取器)适用于多程序工作负载,因为不同的程序映射到不同的内核并独立运行。然而,并行应用程序提出了一系列不同的挑战。并行应用程序的多个线程彼此共享数据,这带来了一致性问题。此外,本地预取器不了解未命中跨线程的不规则分布。在本文中,我们提出了一个硬件预取框架的L1 D-Cache的目标并行应用程序。我们展示了如何通过研究和分类所有线程的L1未命中模式来对L2缓存进行有效的预取请求。我们的初步结果显示,在PARSEC基准测试套件上,平均执行时间提高了7%。
Hardware prefetching has been studied in the past for multiprogrammed workloads as well as GPUs. Efficient hardware prefetchers like stream-based or GHB-based ones work well for multiprogrammed workloads because different programs get mapped to different cores and are run independently. Parallel applications, however, pose a different set of challenges. Multiple threads of a parallel application share data with each other which brings in coherency issues. Also, local prefetchers do not understand the irregular spread of misses across threads. In this paper, we propose a hardware prefetching framework for L1 D-Cache that targets parallel applications. We show how to make efficient prefetch requests to the L2 cache by studying and classifying the patterns of L1 misses across all the threads. Our preliminary results show an improvement of 7% in execution time on an average on the PARSEC benchmark suite.