SB-Fetch: synchronization aware hardware prefetching for chip multiprocessors

SB-Fetch: synchronization aware hardware prefetching for chip multiprocessors
复制标题

SB-Fetch:芯片多处理器的同步感知硬件预取

DOI:
10.1145/3392717.3392735
复制
发表时间:
2020
期刊:
ICS '20: Proceedings of the 34th ACM International Conference on Supercomputing
影响因子:
--
通讯作者:
Jiménez, Daniel A.
Jiménez, Daniel A.
中科院分区:
--
文献类型:
--
作者:
AlBarakat, Laith M.;Gratz, Paul V.;Jiménez, Daniel A.

文献摘要

参考文献

被引文献

相似文献

共享内存的多线程应用程序通常需要程序员在关键部分插入线程同步原语(即锁,屏障和条件变量),以同步进程之间的数据访问。扩展性能需要平衡每个线程的工作负载,在关键部分花费很少的时间。然而,在实践中,线程经常浪费时间等待获取锁/屏障,导致线程不平衡和性能伸缩性差。此外,临界区往往失速的数据预取,以减轻等待的影响,通过确保数据被预加载到核心caches. This时,临界区done. The本文介绍了一种纯硬件技术,使安全的数据预取超过片上多处理器(CMP)的同步点。我们表明,成功的预取超过同步点需要克服现有技术中的两个重大挑战。首先,典型的预取器被设计为基于当前未命中来触发预取。与单线程应用程序中的核心不同,同步点上的多线程核心停顿不会产生新的引用来触发预取器。第二,即使预取被正确地定向为读取超过同步点,它也可能在该数据被写入之前从另一个核预取共享数据。这种预取将被认为是“准确的”,但非常不可取,因为它将由于相干性而导致三个额外的“乒乓”移动,比没有预取花费更多的延迟和能量。我们开发了一个新的数据预取,同步感知B-取(SB-取),内置作为一个扩展到以前的单线程数据预取。SB-Fetch解决了共享内存多线程工作负载的这两个问题。SB-Fetch的新奇在于,它显式地为同步点之外的数据发出预取,并且区分可能和不可能引起缓存一致性开销的数据。这两个特征是直接协同的,因为盲目地预取超过同步可能会导致一致性损失。SB-Fetch使用Parsec [4],Rodinia [7]和Parboil [39]中的一组代表性基准进行评估。SB-Fetch将执行时间比基线提高12.3%,比最佳预取提高4%。
Shared-memory, multi-threaded applications often require programmers to insert thread synchronization primitives (i.e.locks, barriers, and condition variables) in critical sections to synchronize data access between processes. Scaling performance requires balanced per-thread workloads with little time spent in critical sections. In practice, however, threads often waste time waiting to acquire locks/barriers, leading to thread imbalance and poor performance scaling. Moreover, critical sections often stall data prefetchers that mitigate the effects of waiting by ensuring data is preloaded in core caches when the critical section is done.This paper introduces a pure hardware technique to enable safe data prefetching beyond synchronization points in chip multiprocessors (CMPs). We show that successful prefetching beyond synchronization points requires overcoming two significant challenges in existing techniques. First, typical prefetchers are designed to trigger prefetches based on current misses. Unlike cores in single-threaded applications, a multi-threaded core stall on a synchronization point does not produce new references to trigger a prefetcher. Second, even if a prefetch were correctly directed to read beyond a synchronization point, it will likely prefetch shared data from another core before this data has been written. This prefetch would be considered "accurate" but highly undesirable because it would lead to three extra "ping-pong" movements due to coherence, costing more latency and energy than without prefetching. We develop a new data prefetcher, Synchronization-aware B-Fetch (SB-Fetch), built as an extension to a previous single-threaded data prefetcher. SB-Fetch addresses both issues for shared memory multi-threaded workloads. The novelty in SB-Fetch is that it explicitly issues prefetches for data beyond synchronization points and it distinguishes between data likely and unlikely to incur cache coherence overhead. These two features are directly synergistic since blindly prefetching beyond synchronization is likely to incur coherence penalties. No prior work includes both features.SB-Fetch is evaluated using a representative set of benchmarks from Parsec [4], Rodinia [7], and Parboil [39]. SB-Fetch improves execution time by 12.3% over baseline and 4% over best of class prefetching.
DOI: 10.1109/micro.2016.7783763
发表时间: 2016-10
期刊: 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子: --
作者:
Jinchun Kim;Seth H. Pugsley;Paul V. Gratz;A. Reddy;C. Wilkerson;Zeshan A. Chishti
通讯作者: Jinchun Kim;Seth H. Pugsley;Paul V. Gratz;A. Reddy;C. Wilkerson;Zeshan A. Chishti
MTB-Fetch:芯片多处理器的多线程感知硬件预取
DOI: 10.1109/lca.2018.2847345
发表时间: 2018
影响因子: 2.3
作者:
AlBarakat, Laith M.;Gratz, Paul V.;Jimenez, Daniel A.
通讯作者: Jimenez, Daniel A.
适用于新兴并行应用的硬件预取器
DOI: --
发表时间: 2012
期刊: International Conference on Parallel Architectures and Compilation Techniques
影响因子: --
作者:
Biswabandan Panda;S. Balachandran
通讯作者: S. Balachandran
B-Fetch:芯片多处理器的分支预测定向预取
DOI: 10.1109/micro.2014.29
发表时间: 2014
期刊: 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子: --
作者:
David Kadjo;Jinchun Kim;Prabal Sharma;Reena Panda;Paul V. Gratz;Daniel A. Jiménez
通讯作者: Daniel A. Jiménez
多核处理器中软件管理一致性的案例
DOI: --
发表时间: 2010
期刊:
影响因子: --
作者:
Xiaochen Zhou;Hu Chen;Sai Luo;Ying Gao;Shoumeng Yan;B. Lewis
通讯作者: B. Lewis