Collaborative Research: SI2-SSI: A Comprehensive Performance Tuning Framework for the MPI Stack
Collaborative Research: SI2-SSI: A Comprehensive Performance Tuning Framework for the MPI Stack
批准号:
1148424
负责人:
William Barth
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-06-01 至 2016-05-31
中文摘要
消息传递接口(MPI)是现代高端计算(HEC)系统上非常广泛使用的并行编程模型。MPI库的许多性能方面,如延迟、带宽、可扩展性、内存占用、缓存污染、计算和通信的重叠等,都高度依赖于系统配置和应用程序要求。此外,随着多核处理器和商用网络技术(如InfiniBand和10 GigE/iWARP)的发展,现代集群正在迅速变化。它们变得多样化和异构,具有不同数量的处理器内核、处理器速度、存储器速度、多代网络适配器/交换机、I/O接口技术和加速器(GPGPU)等。通常,任何MPI库都通过采用各种运行时参数来处理上述平台多样性和应用程序敏感性。这些参数在发布过程中,或者由系统管理员,或者由最终用户进行调整。 这些默认参数可能对所有系统配置和应用程序都是最佳的,也可能不是最佳的。典型专有系统的MPI库需要为一系列应用程序进行大量的性能调优。 由于商用集群在配置(处理器、内存和网络)方面提供了更大的灵活性,因此使用任何MPI库的发布版本及其默认设置都很难实现最佳调优。这导致了以下广泛的挑战:“能否为MPI库设计一个全面的性能调优框架,以便下一代InfiniBand、10 GigE/iWARP和RoCE集群和应用程序能够提取'裸金属'性能和最大的可扩展性?“调查人员将利用创新的解决方案应对上述挑战,其中包括来自俄亥俄州州立大学(OSU)和俄亥俄州超级计算机中心(OSC)的计算机科学家,以及来自德克萨斯州高级计算中心(TACC)和圣地亚哥超级计算机中心(SDSC)、加州圣地亚哥大学(UCSD)的计算科学家。调查人员将具体应对以下挑战:1)可以设计一组静态工具来优化MPI库在安装时的性能吗? 2)能否设计一组低开销的动态工具来在生产运行期间优化每个用户和每个应用程序的性能? 3)如何将建议的性能调优框架与即将到来的MPIT接口结合起来? 4)如何在给定的系统上配置MPI库,以便为一组驱动应用程序提供不同的优化? 和5)提出的调优框架可以实现哪些好处(在性能、可扩展性、内存效率和减少缓存污染方面)? 这项研究将由NSF计算科学研究人员在TACC Ranger和OSC、SDSC和OSU的其他系统上运行大规模模拟的一系列应用程序驱动。 拟议的设计将被集成到开源MVAPICH 2库中。
英文摘要
The Message Passing Interface (MPI) is a very widely used parallel programming model on modern High-End Computing (HEC) systems. Many performance aspects of MPI libraries, such as latency, bandwidth, scalability, memory footprint, cache pollution, overlap of computation and communication etc. are highly dependent on system configuration and application requirements. Additionally, modern clusters are changing rapidly with the growth of multi-core processors and commodity networking technologies such as InfiniBand and 10GigE/iWARP. They are becoming diverse and heterogeneous with varying number of processor cores, processor speed, memory speed, multi-generation network adapters/switches, I/O interface technologies, and accelerators (GPGPUs), etc. Typically, any MPI library deals with the above kind of diversity in platforms and sensitivity of applications by employing various runtime parameters. These parameters are tuned during its release, or bysystem administrators, or by end-users. These default parameters may or may not be optimal for all system configurations and applications.The MPI library of a typical proprietary system goes through heavy performance tuning for a range of applications. Since commodity clusters provide greater flexibility in their configurations (processor, memory and network), it is very hard to achieve optimal tuning using released version of any MPI library, with its default settings. This leads to the following broad challenge: "Can a comprehensive performance tuning framework be designed for MPI library so that the next generation InfiniBand, 10GigE/iWARP and RoCE clusters and applications will be able to extract `bare-metal' performance and maximum scalability?" The investigators, involving computerscientists from The Ohio State University (OSU) and Ohio Supercomputer Center (OSC) as well as computational scientists from the Texas Advanced Computing Center (TACC) and San Diego Supercomputer Center (SDSC), University of California San Diego (UCSD), will be addressing the above challenge with innovative solutions.The investigators will specifically address the following challenges: 1) Can a set of static tools be designed to optimize performance of an MPI library during installation time? 2) Can a set of dynamic tools with low overhead be designed to optimize performance on a per-user and per-application basis during production runs? 3) How to incorporate the proposed performance tuning framework with the upcoming MPIT interface? 4) How to configure MPI libraries on a given system to deliver different optimizations to a set of driving applications? and 5) What kind of benefits (in terms of performance, scalability, memory efficiency and reduction in cache pollution) can be achieved by the proposed tuning framework? The research will be driven by a set of applications from established NSF computational science researchers running large scale simulations on the TACC Ranger and other systems at OSC, SDSC and OSU. The proposed designs will be integrated into the open-source MVAPICH2 library.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Frameworks: Designing Next-Generation MPI Libraries for Emerging Dense GPU Systems
-
批准号:1931354
-
项目类别:Standard Grant
-
资助金额:$38.32万
-
财政年份:2019
-
负责人:William Barth
-
依托单位:
SHF: Large: Collaborative Research: Next Generation Communication Mechanisms exploiting Heterogeneity, Hierarchy and Concurrency for Emerging HPC Systems
-
批准号:1565431
-
项目类别:Standard Grant
-
资助金额:$42.25万
-
财政年份:2016
-
负责人:William Barth
-
依托单位:
Collaborative Research: Integrated HPC Systems Usage and Performance of Resources Monitoring and Modeling (SUPReMM)
-
批准号:1203604
-
项目类别:Standard Grant
-
资助金额:$45.79万
-
财政年份:2012
-
负责人:William Barth
-
依托单位:
SHF:Large:Collaborative Research:Unified Runtime for Supporting Hybrid Programming Models on Heterogeneous Architecture
-
批准号:1213057
-
项目类别:Standard Grant
-
资助金额:$37.19万
-
财政年份:2012
-
负责人:William Barth
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: