Dense Linear Algebra on Distributed Heterogeneous Hardware with a Symbolic DAG Approach

Dense Linear Algebra on Distributed Heterogeneous Hardware with a Symbolic DAG Approach
复制标题

采用符号 DAG 方法的分布式异构硬件上的密集线性代数

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
G. Bosilca
G. Bosilca
中科院分区:
--
文献类型:
--
作者:
G. Bosilca

文献摘要

被引文献

相似文献

分布式异构硬件上的稠密线性代数与符号化DAG方法乔治博西尔卡托马斯埃罗奥雷连布泰勒皮奥特Luszczek安东尼Danalis杰克J. Dongarra 2012年1月24日简介和动机在微处理器和高端系统设计中发生重大变化的各种因素中,有三个特别值得注意:1.每个芯片的晶体管数量将继续当前的趋势,即大约每18个月翻一番,而处理器时钟的速度将停止增加; 2.对CPU引脚数量和带宽的物理限制正在成为近期的现实; 3.对于千万亿次(和更大的)系统,正在发生向混合/异构系统的强烈漂移。虽然前两个涉及基本的物理限制,目前的技术趋势不太可能在短期内克服,但第三个是前两个的明显后果,加上使用数千个计算单元来扩展到千万亿次和更大的系统的经济必要性。更多的晶体管和更慢的时钟需要多核设计和增加的并行性。传统处理器设计的基本法则--增加晶体管密度、加快时钟速率、降低电压--现在已经被一系列物理障碍所阻止:产生过多的热量、消耗太多的功率、泄漏太多的能量、有用的信号被噪声所克服。多核设计是对这种情况的自然进化反应。通过将多个处理器核心放在一个芯片上,架构师可以克服以前的限制,并继续增加每个芯片的门数,而不增加功率密度。然而,由于过量的热量产生意味着频率不能进一步增加,随着浅宽管道设计成为常态,深窄管道模型将趋于消退。此外,尽管有明显的相似之处,但多核处理器并不等同于多CPU或SMP。同一芯片上的多个内核可以共享各种
Dense Linear Algebra on Distributed Heterogeneous Hardware with a Symbolic DAG Approach George Bosilca Thomas Herault Aurelien Bouteiller Piotr Luszczek Anthony Danalis Jack J. Dongarra January 24, 2012 Introduction and Motivation Among the various factors that drive the momentous changes occurring in the design of microprocessors and high end systems [1], three stand out as especially notable: 1. the number of transistors per chip will continue the current trend, i.e. double roughly every 18 months, while the speed of processor clocks will cease to in- crease; 2. the physical limit on the number and bandwidth of the CPUs pins is becoming a near-term reality; 3. a strong drift toward hybrid/heterogeneous systems for petascale (and larger) systems is taking place. While the first two involve fundamental physical limitations that current technology trends are unlikely to overcome in the near term, the third is an obvious consequence of the first two, combined with the economic necessity of using many thousands of computational units to scale up to petascale and larger systems. More transistors and slower clocks require multicore designs and an increased par- allelism. The fundamental laws of traditional processor design – increasing transistor density, speeding up clock rate, lowering voltage – have now been stopped by a set of physical barriers: excess heat produced, too much power consumed, too much energy leaked, and useful signal overcome by noise. Multicore designs are a natural evolu- tionary response to this situation. By putting multiple processor cores on a single die, architects can overcome the previous limitations, and continue to increase the num- ber of gates per chip without increasing the power densities. However, since excess heat production means that frequencies cannot be further increased, deep-and-narrow pipeline models will tend to recede as shallow-and-wide pipeline designs become the norm. Moreover, despite obvious similarities, multicore processors are not equiva- lent to multiple-CPUs or to SMPs. Multiple cores on the same chip can share various