Throughput computing

Throughput computing
复制标题

吞吐量计算

DOI:
10.1145/1810085.1810088
复制
发表时间:
2010
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
W. Dally
W. Dally
中科院分区:
--
文献类型:
--
作者:
W. Dally

文献摘要

被引文献

相似文献

半导体技术规模的质变已经结束了过去十年来用作高性能计算机构建块的单线程处理器的性能扩展,并使各种规模的计算机的功率受到限制。在当今功率有限的情况下,高效的高性能计算机必须由吞吐量处理器构建,这些处理器(例如 GPU)针对每单位功率的持续性能进行了优化,而不是单线程性能。本演讲将讨论未来吞吐量处理器的架构和编程中的一些挑战和机遇。在这些处理器中,性能源自并行性,效率源自局部性。并行性可以利用吞吐量处理器中丰富且廉价的算术单元。然而,如果没有局部性,带宽很快就会成为瓶颈。现代计算系统中决定成本、性能和功耗的关键资源是通信带宽,而不是算术。本演讲将通过来自 Imagine 和 Merrimac 项目、NVIDIA GPU 以及三代流编程系统的示例来讨论并行性和局部性的利用。
A qualitative change in the scaling of semiconductor technology has ended the performance scaling of the single-thread processors that have been used as the building blocks for high-performance computers for the last decade and has made computers of all scales power limited. In today's power-limited regime, efficient high-performance computers must be built from throughput processors, processors, like GPUs, that are optimized for sustained performance per unit power --- rather than for single-thread performance. This talk will discuss some of the challenges and opportunities in the architecture and programming of future throughput processors. In these processors, performance derives from parallelism and efficiency derives from locality. Parallelism can take advantage of the plentiful and inexpensive arithmetic units in a throughput processor. Without locality, however, bandwidth quickly becomes a bottleneck. Communication bandwidth, not arithmetic is the critical resource in a modern computing system that dominates cost, performance, and power. This talk will discuss exploitation of parallelism and locality with examples drawn from the Imagine and Merrimac projects, from NVIDIA GPUs, and from three generations of stream programming systems.