Optimizing SMT Processors for High Single-Thread Performance

Optimizing SMT Processors for High Single-Thread Performance
复制标题

优化 SMT 处理器以获得高单线程性能

DOI:
--
复制
发表时间:
2003
期刊:
J. Instr. Level Parallelism
影响因子:
--
通讯作者:
Seungryul Choi
Seungryul Choi
中科院分区:
--
文献类型:
--
作者:
Gautham Thambidorai;D. Yeung;Seungryul Choi

文献摘要

被引文献

相似文献

同时多线程(SMT)处理器以牺牲单线程性能为代价实现高处理器吞吐量。本文研究了SMT处理器的资源分配策略,尽可能多地保留指定的“前台”线程的单线程性能,同时仍然允许其他“后台”线程共享资源。由于SMT机器上的后台线程对前台线程的性能影响为零,因此我们将后台线程称为透明线程。透明线程是执行低优先级或非关键计算的理想选择,应用程序可以进行进程调度、从属多线程和在线性能监视。为了实现透明线程,我们提出了三种机制来保持后台线程的透明度:插槽优先级,后台线程优先窗口分区,和后台线程刷新。此外,我们提出了三种机制,以提高后台线程的性能,而不牺牲透明度:积极的提取分区,前台线程执行窗口分区,和前台线程刷新。我们实现了我们的机制上的SMT处理器的详细模拟器,并使用8个基准测试,包括7个从SPEC CPU2000套件进行评估。我们的研究结果表明,当缓存和分支预测干扰的因素,后台线程引入不到1%的性能下降的前景线程。此外,保持后台线程的透明性,相对于同等优先级方案,仅将其吞吐量降低了23%。为了证明透明线程的有用性,我们研究了透明软件预取(TSP),一种使用透明线程的软件数据预取的实现。由于其接近零的开销,TSP可以为程序中的所有加载启用预取检测,从而消除了分析的需要。TSP,没有任何配置文件信息,实现了9.41%的增益在6个SPEC基准,而传统的软件预取引导缓存未命中配置文件的性能仅提高了2.47%。
Simultaneous Multithreading (SMT) processors achieve high processor throughput at the expense of single-thread performance. This paper investigates resource allocation policies for SMT processors that preserve, as much as possible, the single-thread performance of designated “foreground” threads, while still permitting other “background” threads to share resources. Since background threads o ns uch an SMT machine have an ear-zero performance impact on foreground threads, we refer to the background threads as transparent threads. Transparent threads are ideal for performing low-priority or non-critical computations, with applications in process scheduling, subordinate multithreading, and on-line performance monitoring. To realize transparent threads, we propose three mechanisms for maintaining the transparency of background threads: slot prioritization, background thread instruction-window partitioning, and background thread flushing. In addition, we propose three mechanisms to boost background thread performance without sacrificing transparency: aggressive fetch partitioning, foreground thread instruction-window partitioning, and foreground thread flushing. We implement our mechanisms on a detailed simulator of an SMT processor, and evaluate them using 8 benchmarks, including 7 from the SPEC CPU2000 suite. Our results show when cache and branch predictor interference are factored out, background threads introduce less than 1% performance degradation on the foreground thread. Furthermore, maintaining the transparency of background threads reduces their throughput by only 23% relative to an equal priority scheme. To demonstrate the usefulness of transparent threads, we study Transparent Software Prefetching (TSP), an implementation of software data prefetching using transparent threads. Due to its near-zero overhead, TSP enables prefetch instrumentation for all loads in a program, eliminating the need for profiling. TSP, without any profile information, achieves a 9.41% gain across 6 SPEC benchmarks, whereas conventional software prefetching guided by cache-miss profiles increases performance by only 2.47%.