Computation for the Endless Frontier
Computation for the Endless Frontier
批准号:
1818253
负责人:
Daniel Stanzione
金额:
$6000.0万
依托单位国家:
美国
项目类别:
Cooperative Agreement
财政年份:
2018
资助国家:
美国
项目状态:
未结题
起止时间:
2018-09-01 至 2025-02-28
中文摘要
计算对我们国家在科学和工程方面的进步至关重要。无论是通过模拟实验成本高昂或不可能的现象,通过大规模数据分析筛选科学仪器可以产生的海量数字数据,还是通过机器学习从这些海量数据中找到模式并提出假设,计算都是几乎每个科学和工程领域都依赖的通用工具,以加速它们的进步。该项目将部署一个功能强大的新系统,称为“前沿”,它建立在设计理念和操作方法的基础上,德克萨斯州先进计算中心(TACC)在提供领先的计算科学仪器方面的成功证明了这一点。FronTier在NSF的网络基础设施中提供了一个前所未有的规模的系统,该系统将在第一天产生生产性科学,同时也为未来向能力更强的系统的转变做好准备。FronTier是传统中央处理器(CPU)和图形处理器(GPU)的混合系统,其性能远远超过NSF之前的领先级别的计算投资。重要的是,Frontier的设计将支持当前NSF领导级计算应用程序向新系统的无缝过渡,并支持未来预期的新的大规模数据密集型和机器学习工作负载。在部署之后,该项目将与十个学术合作伙伴合作运行该系统。此外,该项目将开始与来自全国各地的领先计算科学家和技术人员合作规划活动,并将利用战略公私合作伙伴关系设计一个领先的计算设施,其科学和工程研究的性能至少要高出十倍,以确保我们国家的整体经济竞争力和繁荣。TACC将与戴尔EMC和英特尔合作,部署Frontier,这是一种混合系统,提供39 pF(双精度)的Intel Xeon处理器,以及11 pF(单精度)的用于机器学习应用的GPU卡。除了NSF以前领先的计算系统主计算节点的3倍每节点内存外,Frontier还将在存储层次结构中拥有2倍的存储带宽,其中包括55PB的可用磁盘存储和3PB的全闪存存储,以支持下一代数据密集型应用程序和对数据科学界的支持。FronTier将部署在TACC最先进的数据中心,该中心配置为从可再生能源中提供系统电力需求的30%。FronTier将通过其对应用程序容器的软件环境支持,以及通过与十个学术机构的伙伴关系,为几乎所有学科的科学和工程提供支持,这些机构提供深厚的计算科学专业知识,为系统用户提供支持。性能至少提高10倍的第二阶段系统的项目规划工作将纳入社区驱动的过程,其中将包括来自全国各地的领先计算科学家和技术专家,并利用战略性的公私合作伙伴关系。这一过程将确保未来NSF领导级计算设施的设计包含最具生产力的短期技术,并预测需要领导级计算和数据分析能力的所有科学和工程领域最有可能的未来技术能力。此外,该项目预计将为领导级计算和数据驱动应用程序开发新的专业知识和技术,将通过出版物、培训和咨询使世界各地的未来用户受益。该项目将利用该团队在教育、外展和培训活动方面的独特方法来鼓励、教育和发展下一代领导力级别的计算科学研究人员。该团队包括校园桥梁、少数群体服务研究所(MSI)外展和数据技术方面的领导者,他们将监督使用Frontier为传统和数据驱动的应用程序使用领导级计算来增加群体多样性的努力。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Computation is critical to our nation's progress in science and engineering. Whether through simulation of phenomena where experiments are costly or impossible, large scale data analysis to sift the enormous quantities of digital data scientific instruments can produce, or machine learning to find patterns and suggest hypothesis from this vast array of data, computation is the universal tool upon which nearly every field of science and engineering relies upon to hasten their advance. This project will deploy a powerful new system, called "Frontier", that builds upon a design philosophy and operations approach proven by the success of the Texas Advanced Computing Center (TACC) in delivering leading instruments for computational science. Frontier provides a system of unprecedented scale in the NSF cyberinfrastructure that will yield productive science on day one, while also preparing the research community for the shift to much more capable systems in the future. Frontier is a hybrid system of conventional Central Processing Units (CPU) and Graphics Processing Units (GPU), with performance capabilities that significantly exceeds prior leadership-class computing investments made by NSF. Importantly, the design of Frontier will support the seamless transition of current NSF leadership-class computing applications to the new system, as well as enable new large-scale data-intensive and machine learning workloads that are expected in the future. Following deployment, the project will operate the system in partnership with ten academic partners. In addition, the project will begin planning activities in collaboration with leading computational scientists and technologists from around the country, and will leverage strategic public-private partnerships to design a leadership-class computing facility with at least ten times more performance capabilities for Science and Engineering research, ensuring the economic competitiveness and prosperity for our nation at large.TACC, in partnerships with Dell EMC and Intel, will deploy Frontier, a hybrid system offering 39 PF (double precision) of Intel Xeon processors, complemented by 11 PF (single precision) of GPU cards for machine learning applications. In addition to 3x the per node memory of NSF's prior leadership-class computing system primary compute nodes, Frontier will have 2x the storage bandwidth in a storage hierarchy that includes 55PB of usable disk-based storage and 3PB of 'all flash' storage, to enable next generation data-intensive applications and support for the data science community. Frontier will be deployed in TACC's state-of-the-art datacenter which is configured to supply 30% of the system's power needs from renewable energy. Frontier will include support for science and engineering in virtually all disciplines through its software environment support for application containers, as well as through its partnership with ten academic institutions providing deep computational science expertise in support of users on the system. The project planning effort for a Phase 2 system with at least 10x performance improvement will incorporate a community-driven process that will include leading computational scientists and technologists from around the country and leverage strategic public-private partnerships. This process will ensure the design of a future NSF leadership-class computing facility that incorporates the most productive near-term technologies, and anticipates the most likely future technological capabilities for all of science and engineering requiring leadership-class computational and data-analytics capabilities. Furthermore, the project is expected to develop new expertise and techniques for leadership-class computing and data-driven applications that will benefit future users worldwide through publications, training, and consulting. The project will leverage the team's unique approach to education, outreach, and training activities to encourage, educate, and develop the next generation of leadership-class computational science researchers. The team includes leaders in campus bridging, minority-serving institute (MSI) outreach, and data technologies who will oversee efforts to use Frontier to increase the diversity of groups using leadership-class computing for traditional and data-driven applications.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(12)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3547276.3548524
发表时间:
2022-08
期刊:
Workshop Proceedings of the 51st International Conference on Parallel Processing
影响因子:
--
作者:
[Tu Tran;Benjamin Michalowicz;B. Ramesh;H. Subramoni;A. Shafi;D. Panda]
通讯作者:
Tu Tran;Benjamin Michalowicz;B. Ramesh;H. Subramoni;A. Shafi;D. Panda
Network-Assisted Noncontiguous Transfers for GPU-Aware MPI Libraries
GPU 感知 MPI 库的网络辅助非连续传输
DOI:
10.1109/mm.2023.3241133
发表时间:
2023
期刊:
IEEE Micro
影响因子:
3.6
作者:
[Suresh, Kaushik Kandadi, Khorassani, Kawthar Shafie, Chen, Chen Chun, Ramesh, Bharath, Abduljabbar, Mustafa, Shafi, Aamir, Subramoni, Hari, Panda, Dhabaleswar K.]
通讯作者:
Panda, Dhabaleswar K.
OMB-Py: Python Micro-Benchmarks for Evaluating Performance of MPI Libraries on HPC Systems
OMB-Py:用于评估 HPC 系统上 MPI 库性能的 Python 微基准
DOI:
10.1109/ipdpsw55747.2022.00143
发表时间:
2022
期刊:
23rd Parallel and Distributed Scientific and Engineering Computing Workshop (PDSEC
影响因子:
--
作者:
[Alnaasan, Nawras, Jain, Arpan, Shafi, Aamir, Subramoni, Hari, Panda, Dhabaleswar K]
通讯作者:
Panda, Dhabaleswar K
Hy-Fi: Hybrid Five-Dimensional Parallel DNN Training on High-Performance GPU Clusters
Hy-Fi:高性能 GPU 集群上的混合五维并行 DNN 训练
DOI:
10.1007/978-3-031-07312-0_6
发表时间:
2022
期刊:
Proceedings International Conference on High Performance Computing
影响因子:
--
作者:
[Jain, A, Shafi, A., Anthony, Q., Kousha, P., Subramoni, H., Panda, DK.]
通讯作者:
Panda, DK.
Highly Efficient Alltoall and Alltoallv Communication Algorithms for GPU Systems
适用于 GPU 系统的高效 Alltoall 和 Alltoallv 通信算法
DOI:
10.1109/ipdpsw55747.2022.00014
发表时间:
2022
期刊:
Heterogeneity in Computing Workshop
影响因子:
--
作者:
[Chen, Chen-Chun, Khorassani, Kawthar Shafie, Anthony, Quentin G., Shafi, Aamir, Subramoni, Hari, Panda, Dhabaleswar K.]
通讯作者:
Panda, Dhabaleswar K.
共 12 条
Final Design Planning for the Leadership-Class Computing Facility
-
批准号:2212090
-
项目类别:Cooperative Agreement
-
资助金额:$350.0万
-
财政年份:2022
-
负责人:Daniel Stanzione
-
依托单位:
Characteristic Science Applications for the Leadership Class Computing Facility
-
批准号:2139536
-
项目类别:Cooperative Agreement
-
资助金额:$699.94万
-
财政年份:2021
-
负责人:Daniel Stanzione
-
依托单位:
Preliminary Design Planning for the Leadership-Class Computing Facility
-
批准号:2033468
-
项目类别:Cooperative Agreement
-
资助金额:$350.0万
-
财政年份:2020
-
负责人:Daniel Stanzione
-
依托单位:
Collaborative Research: Chameleon Phase III: A Large-Scale, Reconfigurable Experimental Environment for Cloud Research
-
批准号:2027176
-
项目类别:Cooperative Agreement
-
资助金额:$300.1万
-
财政年份:2020
-
负责人:Daniel Stanzione
-
依托单位:
Planning for the Leadership-Class Computing Facility
-
批准号:1925096
-
项目类别:Cooperative Agreement
-
资助金额:$200.0万
-
财政年份:2019
-
负责人:Daniel Stanzione
-
依托单位:
Planning for the Leadership-Class Computing Facility
-
批准号:1940979
-
项目类别:Cooperative Agreement
-
资助金额:$0.0万
-
财政年份:2019
-
负责人:Daniel Stanzione
-
依托单位:
Operations & Maintenance for the Endless Frontier
-
批准号:1854828
-
项目类别:Cooperative Agreement
-
资助金额:$6000.0万
-
财政年份:2019
-
负责人:Daniel Stanzione
-
依托单位:
Stampede 2: Operations and Maintenance for the Next Generation of Petascale Computing
-
批准号:1663578
-
项目类别:Cooperative Agreement
-
资助金额:$2400.0万
-
财政年份:2017
-
负责人:Daniel Stanzione
-
依托单位:
Collaborative Research: Chameleon: A Large-Scale, Reconfigurable Experimental Environment for Cloud Research
-
批准号:1743354
-
项目类别:Cooperative Agreement
-
资助金额:$353.92万
-
财政年份:2017
-
负责人:Daniel Stanzione
-
依托单位:
Stampede 2: The Next Generation of Petascale Computing for Science and Engineering
-
批准号:1540931
-
项目类别:Cooperative Agreement
-
资助金额:$3000.0万
-
财政年份:2016
-
负责人:Daniel Stanzione
-
依托单位:
Collaborative Research: Chameleon: A Large-Scale, Reconfigurable Experimental Environment for Cloud Research
-
批准号:1419152
-
项目类别:Cooperative Agreement
-
资助金额:$571.35万
-
财政年份:2014
-
负责人:Daniel Stanzione
-
依托单位:
Wrangler: A Transformational Data Intensive Resource for the Open Science Community
-
批准号:1341711
-
项目类别:Cooperative Agreement
-
资助金额:$600.0万
-
财政年份:2013
-
负责人:Daniel Stanzione
-
依托单位:
Collaborative Research: The Science Gateway Institute (SGW-I) for the Democratization and Acceleration of Science
-
批准号:1216733
-
项目类别:Standard Grant
-
资助金额:$6.5万
-
财政年份:2012
-
负责人:Daniel Stanzione
-
依托单位:
Enabling, Enhancing, and Extending Petascale Computing for Science and Engineering
-
批准号:1134872
-
项目类别:Cooperative Agreement
-
资助金额:$2750.0万
-
财政年份:2011
-
负责人:Daniel Stanzione
-
依托单位:
Increasing Student Participation in Cluster Computing through IEEE Cluster 2011 Attendance
-
批准号:1152113
-
项目类别:Standard Grant
-
资助金额:$3.5万
-
财政年份:2011
-
负责人:Daniel Stanzione
-
依托单位:
GDBase: An Engine for Scalable offline Debugging
-
批准号:1019055
-
项目类别:Standard Grant
-
资助金额:$30.21万
-
财政年份:2009
-
负责人:Daniel Stanzione
-
依托单位:
GDBase: An Engine for Scalable offline Debugging
-
批准号:0850853
-
项目类别:Standard Grant
-
资助金额:$30.21万
-
财政年份:2009
-
负责人:Daniel Stanzione
-
依托单位:
海外基金