OUTRIDER: Efficient memory latency tolerance with decoupled strands

OUTRIDER: Efficient memory latency tolerance with decoupled strands
复制标题

OUTRIDER:通过解耦链实现高效的内存延迟容忍度

DOI:
10.1145/2000064.2000079
复制
发表时间:
2011
期刊:
2011 38th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Sanjay J. Patel
Sanjay J. Patel
中科院分区:
--
文献类型:
--
作者:
N. Crago;Sanjay J. Patel

文献摘要

被引文献

相似文献

我们提出Outrider,面向吞吐量的处理器,提供内存延迟容忍,以提高高线程工作负载的性能的架构。Out-rider使单个执行线程能够作为多个解耦的指令流呈现给架构,这些指令流将存储器访问指令和存储器消耗指令分开。关键的见解是,通过解耦指令流,处理器流水线可以以类似于乱序设计的方式容忍内存延迟,同时依赖于低复杂度的有序微架构。此外,Outrider可以容忍内存延迟,而不是像现代GPU那样添加更多线程,并减少线程之间共享资源的争用。我们证明,Outrider可以超越单线程核心的23-131%和4路同步多线程核心高达87%的数据并行应用程序中的1024核系统。此外,Outrider实现了这些性能增益,而不会产生额外的硬件线程上下文的开销,这导致与多线程核心相比提高了区域效率。
We present Outrider, an architecture for throughput-oriented processors that provides memory latency tolerance to improve performance on highly threaded workloads. Out-rider enables a single thread of execution to be presented to the architecture as multiple decoupled instruction streams that separate memory-accessing and memory-consuming instructions. The key insight is that by decoupling the instruction streams, the processor pipeline can tolerate memory latency in a way similar to out-of-order designs while relying on a low-complexity in-order micro-architecture. Moreover, instead of adding more threads as is done in modern GPUs, Outrider can tolerate memory latency with fewer threads and reduced contention for resources shared amongst threads. We demonstrate that Outrider can outperform single threaded cores by 23-131% and a 4-way simultaneous multithreaded core by up to 87% on data parallel applications in a 1024-core system. Moreover, Outrider achieves these performance gains without incurring the overhead of additional hardware thread contexts, which results in improved area efficiency compared to a multithreaded core.