SB-Fetch: synchronization aware hardware prefetching for chip multiprocessors
SB-Fetch: synchronization aware hardware prefetching for chip multiprocessors
复制标题
SB-Fetch:芯片多处理器的同步感知硬件预取
DOI:
10.1145/3392717.3392735
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Jiménez, Daniel A.
中科院分区:
文献类型:
--
作者:
AlBarakat, Laith M.;Gratz, Paul V.;Jiménez, Daniel A.
Shared-memory, multi-threaded applications often require programmers to insert thread synchronization primitives (i.e.locks, barriers, and condition variables) in critical sections to synchronize data access between processes. Scaling performance requires balanced per-thread workloads with little time spent in critical sections. In practice, however, threads often waste time waiting to acquire locks/barriers, leading to thread imbalance and poor performance scaling. Moreover, critical sections often stall data prefetchers that mitigate the effects of waiting by ensuring data is preloaded in core caches when the critical section is done.This paper introduces a pure hardware technique to enable safe data prefetching beyond synchronization points in chip multiprocessors (CMPs). We show that successful prefetching beyond synchronization points requires overcoming two significant challenges in existing techniques. First, typical prefetchers are designed to trigger prefetches based on current misses. Unlike cores in single-threaded applications, a multi-threaded core stall on a synchronization point does not produce new references to trigger a prefetcher. Second, even if a prefetch were correctly directed to read beyond a synchronization point, it will likely prefetch shared data from another core before this data has been written. This prefetch would be considered "accurate" but highly undesirable because it would lead to three extra "ping-pong" movements due to coherence, costing more latency and energy than without prefetching. We develop a new data prefetcher, Synchronization-aware B-Fetch (SB-Fetch), built as an extension to a previous single-threaded data prefetcher. SB-Fetch addresses both issues for shared memory multi-threaded workloads. The novelty in SB-Fetch is that it explicitly issues prefetches for data beyond synchronization points and it distinguishes between data likely and unlikely to incur cache coherence overhead. These two features are directly synergistic since blindly prefetching beyond synchronization is likely to incur coherence penalties. No prior work includes both features.SB-Fetch is evaluated using a representative set of benchmarks from Parsec [4], Rodinia [7], and Parboil [39]. SB-Fetch improves execution time by 12.3% over baseline and 4% over best of class prefetching.
登录
查看更多内容
DOI:
10.1109/micro.2016.7783763
发表时间:
2016-10
期刊:
2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
作者:
Jinchun Kim;Seth H. Pugsley;Paul V. Gratz;A. Reddy;C. Wilkerson;Zeshan A. Chishti
通讯作者:
Jinchun Kim;Seth H. Pugsley;Paul V. Gratz;A. Reddy;C. Wilkerson;Zeshan A. Chishti
影响因子:
2.3
作者:
AlBarakat, Laith M.;Gratz, Paul V.;Jimenez, Daniel A.
通讯作者:
Jimenez, Daniel A.
DOI:
--
发表时间:
2012
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
作者:
Biswabandan Panda;S. Balachandran
通讯作者:
S. Balachandran
DOI:
10.1109/micro.2014.29
发表时间:
2014
期刊:
2014 47th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
作者:
David Kadjo;Jinchun Kim;Prabal Sharma;Reena Panda;Paul V. Gratz;Daniel A. Jiménez
通讯作者:
Daniel A. Jiménez
DOI:
--
发表时间:
2010
期刊:
影响因子:
--
作者:
Xiaochen Zhou;Hu Chen;Sai Luo;Ying Gao;Shoumeng Yan;B. Lewis
通讯作者:
B. Lewis