Hardware prefetchers for emerging parallel applications
Hardware prefetchers for emerging parallel applications
复制标题
适用于新兴并行应用的硬件预取器
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
S. Balachandran
中科院分区:
文献类型:
--
作者:
Biswabandan Panda;S. Balachandran
Hardware prefetching has been studied in the past for multiprogrammed workloads as well as GPUs. Efficient hardware prefetchers like stream-based or GHB-based ones work well for multiprogrammed workloads because different programs get mapped to different cores and are run independently. Parallel applications, however, pose a different set of challenges. Multiple threads of a parallel application share data with each other which brings in coherency issues. Also, local prefetchers do not understand the irregular spread of misses across threads. In this paper, we propose a hardware prefetching framework for L1 D-Cache that targets parallel applications. We show how to make efficient prefetch requests to the L2 cache by studying and classifying the patterns of L1 misses across all the threads. Our preliminary results show an improvement of 7% in execution time on an average on the PARSEC benchmark suite.