SI2-SSI: FAMII: High Performance and Scalable Fabric Analysis, Monitoring and Introspection Infrastructure for HPC and Big Data
SI2-SSI: FAMII: High Performance and Scalable Fabric Analysis, Monitoring and Introspection Infrastructure for HPC and Big Data
批准号:
1664137
负责人:
Dhabaleswar Panda
金额:
$80.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-01 至 2020-06-30
中文摘要
随着计算、网络、异构硬件和存储技术在高端计算 (HEC) 平台中不断发展,了解时间关键的高性能计算 (HPC) 和大数据应用程序之间的交互、它们实现高性能便携式解决方案所依赖的软件基础设施、这些高性能中间件所依赖的底层通信结构以及管理 HPC 集群的调度程序之间的交互变得越来越重要和具有挑战性。 这种理解将使所有相关方(应用程序开发人员/用户、系统管理员和中间件开发人员)能够最大限度地提高组成现代 HPC 系统的各个组件的效率和性能,并解决不同的重大挑战问题。显然需要但不幸的是缺乏一种高性能和可扩展的工具,能够分析结构上的通信并将其与 HPC/大数据应用程序、底层中间件和现有大型 HPC 系统上的作业调度程序的行为关联起来。 拟议的协同和协作工作由来自 OSU 和 OSC 的计算机和计算科学家团队进行,旨在创建一个集成的软件基础设施,用于 HPC 和大数据的高性能和可扩展的结构分析、监控和自省。该工具将实现以下目标:1)可移植、易于使用且易于理解,2)具有高性能和可扩展的渲染和存储技术,3)适用于可能在现有大型 HPC 系统和新兴的百亿亿次系统上使用的不同通信结构和编程模型。 拟议的研究和开发工作的变革性影响是为当前和下一代千万亿级/百亿亿级系统的应用设计一个全面的分析和性能监控工具,以利用最大的性能和可扩展性。拟议的研究和相关基础设施将对实现以前难以提供的高性能计算和大数据应用程序的优化产生重大影响。这些潜在的成果将通过使用拟议的框架来验证多种场景下的各种 HPC 和大数据基准和应用程序来证明。 集成的中间件和工具将通过顶级论坛的公共存储库和出版物向社区公开,从而使其他 MPI 和大数据堆栈能够采用这些设计。 研究结果还将传播给研究人员的合作组织,以影响他们的 HPC 软件产品和应用程序。拟议的研究方向及其解决方案将用于 PI 的课程中,以培训本科生和研究生,包括代表性不足的少数族裔和女学生。该提案解决的技术挑战包括:1)大型复杂 HEC 网络的可扩展可视化,以便为最终用户提供近乎即时的渲染,2)可轻松移植到多种通信结构、新型计算架构和高性能中间件的通用数据收集方案,3)通过优化数据库模式和使用内存支持的键值存储/数据库来增强数据存储性能,4)MPI、PGAS 和大数据库的支持,以实现拟议的监控、分析和内省框架,以及 5) 实现对特定应用领域的更深入内省。 研究还将由一组 HPC 和大数据应用程序驱动。拟议的研发工作的变革性影响是为当前和下一代千万亿级/百亿亿级系统的应用设计全面的分析和性能监控工具,以利用最大的性能和可扩展性。
英文摘要
As the computing, networking, heterogeneous hardware, and storagetechnologies continue to evolve in High-End Computing (HEC) platforms,it becomes increasingly essential and challenging to understand theinteractions between time-critical High-Performance Computing (HPC)and Big Data applications, the software infrastructures upon whichthey rely for achieving high-performing portable solutions, theunderlying communication fabric these high-performance middlewaresdepend on and the schedulers that manage HPC clusters. Suchunderstanding will enable all involved parties (applicationdevelopers/users, system administrators, and middleware developers) tomaximize the efficiency and performance of the individual componentsthat comprise a modern HPC system and solve different grand challengeproblems. There is a clear need and unfortunate lack of a high-performance andscalable tool that is capable of analyzing and correlating thecommunication on the fabric with the behavior of HPC/Big Dataapplications, underlying middleware and the job scheduler on existinglarge HPC systems. The proposed synergistic and collaborative effort,undertaken by a team of computer and computational scientists from OSUand OSC, aims to create an integrated software infrastructure for high-performance and scalable Fabric Analysis, Monitoring andIntrospection for HPC and Big Data. This tool will achieve thefollowing objectives: 1) be portable, easy to use and easy tounderstand, 2) have high performance and scalable rendering andstorage techniques and, 3) be applicable to the differentcommunication fabrics and programming models that are likely to beused on existing large HPC systems and emerging exascale systems. Thetransformative impact of the proposed research and development effortis to design a comprehensive analysis and performance monitoring toolfor applications of current and next generation multipetascale/exascale systems to harness the maximum performance andscalability.The proposed research and the associated infrastructure will have asignificant impact on enabling optimizations of HPC and Big Dataapplications that have previously been difficult to provide. Thesepotential outcomes will be demonstrated by using the proposedframework to validate a variety of HPC and Big Data benchmarks andapplications under multiple scenarios. The integrated middleware andtools will be made publicly available to the community through publicrepositories and publications in the top forums, enabling other MPIand Big Data stacks to adopt the designs. Research results will alsobe disseminated to the collaborating organizations of theinvestigators to impact their HPC software products andapplications. The proposed research directions and their solutionswill be used in the curriculum of the PIs to train undergraduate andgraduate students, including under-represented minorities and femalestudents. The technical challenges addressed by the proposal include: 1)Scalable visualization of large and complex HEC networks so as toprovide a near instant rendering to end users, 2) A generalized datagathering scheme which is easily portable to multiple communicationfabrics, novel compute architectures and high-performance middleware,3) Enhanced data storage performance through optimized databaseschemas and the use of memory-backed key value stores/databases, 4)Support in MPI, PGAS, and Big Data libraries to enable the proposedmonitoring, analysis, and introspection framework, and 5) Enablingdeeper introspection of particular regions of application. Theresearch will also be driven by a set of HPC and Big Dataapplications. The transformative impact of the proposed research anddevelopment effort is to design a comprehensive analysis andperformance monitoring tool for applications of current and nextgeneration multi petascale/exascale systems to harness the maximumperformance and scalability.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
A Scalable Network-Based Performance Analysis Tool for MPI on Large-Scale HPC Systems
大规模 HPC 系统上 MPI 的可扩展的基于网络的性能分析工具
DOI:
10.1109/cluster.2017.78
发表时间:
2017
期刊:
IEEE International Conference on Cluster Computing (CLUSTER
影响因子:
--
作者:
[Subramoni, Hari, Lu, Xiaoyi, Panda, Dhabaleswar K.]
通讯作者:
Panda, Dhabaleswar K.
DOI:
10.1109/ipdps.2019.00034
发表时间:
2019-05
期刊:
2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
作者:
[Jie Zhang;Xiaoyi Lu;Ching-Hsiang Chu;D. Panda]
通讯作者:
Jie Zhang;Xiaoyi Lu;Ching-Hsiang Chu;D. Panda
Designing a Profiling and Visualization Tool for Scalable and In-depth Analysis of High-Performance GPU Clusters
设计用于对高性能 GPU 集群进行可扩展和深入分析的分析和可视化工具
DOI:
10.1109/hipc.2019.00022
发表时间:
2019
期刊:
and Analytics (HiPC
影响因子:
--
作者:
[Kousha, Pouya, Ramesh, Bharath, Kandadi Suresh, Kaushik, Chu, Ching-Hsiang, Jain, Arpan, Sarkauskas, Nick, Subramoni, Hari, Panda, Dhabaleswar K.]
通讯作者:
Panda, Dhabaleswar K.
DOI:
10.1007/978-3-030-32813-9_18
发表时间:
2018-12
期刊:
影响因子:
--
作者:
[Haiyang Shi;Xiaoyi Lu;D. Panda]
通讯作者:
Haiyang Shi;Xiaoyi Lu;D. Panda
Accelerated Real-time Network Monitoring and Profiling at Scale using OSU INAM
使用 OSU INAM 加速实时网络监控和大规模分析
DOI:
10.1145/3311790.3396672
发表时间:
2020
期刊:
2020 PEARC: Practice and Experience in Advanced Research Computing
影响因子:
--
作者:
[Kousha, P., S. D., Kamal Raj, Subramoni, H., Panda, D. K., Na, H., Dockendorf, T., Tomko, K.]
通讯作者:
Tomko, K.
共 6 条
CSR: Small: CONCERT: Designing Scalable Communication Runtimes with On-the-fly Compression for HPC and AI Applications on Heterogeneous Architectures
-
批准号:2312927
-
项目类别:Standard Grant
-
资助金额:$60.0万
-
财政年份:2023
-
负责人:Dhabaleswar Panda
-
依托单位:
Travel: Student Travel Support for MVAPICH User Group (MUG) 2023 Conference
-
批准号:2331223
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2023
-
负责人:Dhabaleswar Panda
-
依托单位:
Collaborative Research: Frameworks: Performance Engineering Scientific Applications with MVAPICH and TAU using Emerging Communication Primitives
-
批准号:2311830
-
项目类别:Standard Grant
-
资助金额:$90.0万
-
财政年份:2023
-
负责人:Dhabaleswar Panda
-
依托单位:
Travel: Student Travel Support for MVAPICH User group (MUG) 2022 Conference
-
批准号:2231825
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2022
-
负责人:Dhabaleswar Panda
-
依托单位:
AI Institute for Intelligent CyberInfrastructure with Computational Learning in the Environment (ICICLE)
-
批准号:2112606
-
项目类别:Cooperative Agreement
-
资助金额:$2000.0万
-
财政年份:2021
-
负责人:Dhabaleswar Panda
-
依托单位:
MRI: RADiCAL: Reconfigurable Major Research Cyberinfrastructure for Advanced Computational Data Analytics and Machine Learning
-
批准号:2018627
-
项目类别:Standard Grant
-
资助金额:$77.0万
-
财政年份:2020
-
负责人:Dhabaleswar Panda
-
依托单位:
OAC Core: Small: Next-Generation Communication and I/O Middleware for HPC and Deep Learning with Smart NICs
-
批准号:2007991
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2020
-
负责人:Dhabaleswar Panda
-
依托单位:
Student Travel Support for MVAPICH User Group (MUG) Meeting
-
批准号:1930003
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2019
-
负责人:Dhabaleswar Panda
-
依托单位:
Collaborative Research: Frameworks: Designing Next-Generation MPI Libraries for Emerging Dense GPU Systems
-
批准号:1931537
-
项目类别:Standard Grant
-
资助金额:$136.05万
-
财政年份:2019
-
负责人:Dhabaleswar Panda
-
依托单位:
Student Travel Support for MVAPICH User Group (MUG) Meeting
-
批准号:1839739
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2018
-
负责人:Dhabaleswar Panda
-
依托单位:
Student Travel Support for MVAPICH User group (MUG) Meeting
-
批准号:1744956
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2017
-
负责人:Dhabaleswar Panda
-
依托单位:
SHF: Large: Collaborative Research: Next Generation Communication Mechanisms exploiting Heterogeneity, Hierarchy and Concurrency for Emerging HPC Systems
-
批准号:1565414
-
项目类别:Standard Grant
-
资助金额:$117.19万
-
财政年份:2016
-
负责人:Dhabaleswar Panda
-
依托单位:
BD Spokes: SPOKE: MIDWEST: Collaborative:Advanced Computational Neuroscience Network (ACNN)
-
批准号:1636846
-
项目类别:Standard Grant
-
资助金额:$16.65万
-
财政年份:2016
-
负责人:Dhabaleswar Panda
-
依托单位:
Student Travel Support for MVAPICH User Group (MUG) Meeting
-
批准号:1644528
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2016
-
负责人:Dhabaleswar Panda
-
依托单位:
SI2-SSI: Collaborative Research: A Software Infrastructure for MPI Performance Engineering: Integrating MVAPICH and TAU via the MPI Tools Interface
-
批准号:1450440
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2015
-
负责人:Dhabaleswar Panda
-
依托单位:
Collaborative Research: Chameleon: A Large-Scale, Reconfigurable Experimental Environment for Cloud Research
-
批准号:1419123
-
项目类别:Cooperative Agreement
-
资助金额:$60.0万
-
财政年份:2014
-
负责人:Dhabaleswar Panda
-
依托单位:
BIGDATA: F: DKM: Collaborative Research: Scalable Middleware for Managing and Processing Big Data on Next Generation HPC Systems
-
批准号:1447804
-
项目类别:Standard Grant
-
资助金额:$72.02万
-
财政年份:2014
-
负责人:Dhabaleswar Panda
-
依托单位:
CSR: Eager: HPC Virtualization with SR-IOV
-
批准号:1347189
-
项目类别:Standard Grant
-
资助金额:$9.83万
-
财政年份:2013
-
负责人:Dhabaleswar Panda
-
依托单位:
Collaborative Research: SI2-SSI: A Comprehensive Performance Tuning Framework for the MPI Stack
-
批准号:1148371
-
项目类别:Standard Grant
-
资助金额:$125.16万
-
财政年份:2012
-
负责人:Dhabaleswar Panda
-
依托单位:
SHF: Large: Collaborative Research: Unified Runtime for Supporting Hybrid Programming Models on Heterogeneous Architecture
-
批准号:1213084
-
项目类别:Standard Grant
-
资助金额:$104.58万
-
财政年份:2012
-
负责人:Dhabaleswar Panda
-
依托单位:
国内基金
海外基金
登录
查看更多内容
考虑SSI效应的导管架式海洋平台抗震性能研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:刘书童
-
依托单位:
考虑SSI的层间隔震高层建筑结构在三维地震下的响应研究
-
批准号:52168072
-
项目类别:地区科学基金项目
-
资助金额:35万元
-
批准年份:2021
-
负责人:刘德稳
-
依托单位:
考虑SSI效应的大型储罐动力学特性及其隔板减晃研究
-
批准号:51978336
-
项目类别:面上项目
-
资助金额:61.0万元
-
批准年份:2019
-
负责人:周叮
-
依托单位:
考虑SSI效应的摇摆墙-框架结构抗震机理及性能评估方法研究
-
批准号:51978524
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2019
-
负责人:李培振
-
依托单位:
考虑能量需求和SSI效应的RC梁式桥基于性能的抗震设计方法
-
批准号:50908014
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2009
-
负责人:江辉
-
依托单位: