Design and Implementation of Algorithms for an Experimental High-Radix Network Switching system
Design and Implementation of Algorithms for an Experimental High-Radix Network Switching system
批准号:
1240652
负责人:
John McCalpin
金额:
$14.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-10-01 至 2013-09-30
中文摘要
计算机硬件的技术和经济趋势已经导致了像今天这样的超大多核服务器集群的广泛采用。超级计算机。当应用程序被修改以利用这些额外节点和核心的增加的并行性时,应用程序?S网络消息通常变得更小、更频繁。尽管在当前的超级计算机系统中,互连网络的带宽几乎可以跟上计算性能的提高,但在发送消息的开销方面几乎没有什么改善,相应地,在发送短消息时,网络的吞吐量也几乎没有提高。由于这两种趋势,许多应用程序在现有的超级计算系统上的可伸缩性很差,而且随着行业向前发展,使用更大的超级计算机系统,预计会有更多应用程序的可伸缩性很差。在这项研究中,TACC正在研究一种新的网络交换机和网络接口架构的可能解决方案,这种架构可以使用非常短的消息来维持整个网络带宽。该项目正在研究有效使用该网络所需的编程模型,并正在评估新互连网络的性能,直接与当前超级计算机中广泛使用的现代(四数据速率Infiniband)互连进行比较。目前正在通过若干个案研究对新制度进行评价。由于网络性能限制,在标准系统上表现出较差的并行扩展算法的实现被移植到新系统上,并使用计时器和硬件性能计数器进行检测,以记录详细的性能特征。这项研究将对这些算法的新架构当前实现的技术可行性进行初步评估,并有望为未来架构的增强提供建议。
英文摘要
Technology and economic trends in computer hardware have led to the widespread adoption of extremely large clusters of multicore servers as today?s supercomputers. As applications are modified to exploit the increased parallelism of these additional nodes and cores, the application?s network messages typically become both smaller and more frequent. Although the bandwidth of the interconnect networks in current supercomputer systems is almost keeping up with increases in compute performance, there has been little improvement in the overhead of sending messages, and correspondingly little improvement in the throughput of the network when sending short messages. As a consequence of these two trends, many applications scale poorly on existing supercomputing systems, and many more applications are expected to scale poorly as the industry moves forward with even larger supercomputer systems. In this study, TACC is investigating a possible solution to these issues with a new network switch and network interface architecture that can sustain full network bandwidth using very short messages. This project is investigating the programming models required to use this network efficiently and is evaluating the performance of the new interconnect network in direct comparison with a modern (quad-data-rate Infiniband) interconnect that is widely used in current supercomputers.The evaluation of the new system is being conducted through a number of case studies. Implementations of algorithms known to exhibit poor parallel scaling on standard systems due to network performance limitations are being ported to the new system and instrumented with timers and hardware performance counters to document the detailed performance characteristics. This study will provide an initial evaluation of the technical viability of the current implementation of this new architecture for these algorithms, and is expected to provide recommendations for future enhancements to the architecture.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Dynamical Balances in a Western Boundary Current: The Effects of Continental Slope/Shelf Topography
-
批准号:9206176
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:1992
-
负责人:John McCalpin
-
依托单位:
海外基金