课题基金 / 基金详情

Elements: Transformation-Based High-Performance Computing in Dynamic Languages

Elements: Transformation-Based High-Performance Computing in Dynamic Languages
要素:动态语言中基于转换的高性能计算
批准号:
1931577
负责人:
Andreas Kloeckner
金额:
$59.97万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-10-01 至 2024-09-30

项目摘要

项目成果

Andreas Kloeckner的其他基金

相似基金

相关文献

中文摘要
翻译
技术计算中的一项关键能力是通过各种不同的过程处理大型规则形状的数字数组。该设施是天气预报、人工智能和图像处理等领域的基础。相应地,现代计算硬件已经进化出了执行这种高效计算的先进能力。不幸的是,到目前为止,将期望的流程调整到给定硬件的过程是昂贵的、费力的,而且容易出错。在一个天真的认识和一个谨慎的认识之间,在表现上有50倍的差异是普遍现象,而不是例外。这个项目的主题Loopy通过使用人工引导的自动程序重写来解决这个问题。从自然和工程现象的模拟到神经科学,Loopy已经在许多应用中展示了应用程序的影响,在这些应用中,它帮助用户以更少的努力获得更高的性能。本提案涉及几个重要的改进,通过扩大Loopy可以转换的程序类别,改进Loopy表示片上通信的手段,并允许它实现通常在有效实施中存在困难的重要基本操作,这些改进将有助于使Loopy更有效和更容易应用。这项工作的一个重要组成部分是通过实现交互式用户界面,使Loopy本身易于其用户社区使用,这样程序转换就可以通过点击鼠标来应用,而不是通过编写计算机代码。所提出的进展将通过一个示例工作负载进行演示,该示例工作负载是当今技术计算中面临的许多计算和软件挑战的象征。多维数组(有时称为“张量”)是许多科学计算的基础数据结构,其应用范围从天气预报到深度学习,再到图像处理和计算神经科学。即使是最简单的数组运算之一——矩阵-矩阵乘法的有效执行,也给现代计算机带来了相当大的技术挑战。通过基于多面体的程序转换工具,所提出的软件将提供数学意图和程序优化的技术挑战之间的分离,允许每个任务由领域专家执行。在提议的项目中,PI将开发更有效的片上通信,前缀和代码生成,程序转换中的重用和抽象,增加转换发现和性能分析的易用性,以及在用户程序中表示数组计算的方法。PI将通过具有广泛适用性的具有挑战性的应用程序验证所建议的技术。所提出的研究的智力价值在于(1)绘制和扩展基于转换的编程的景观,从一次性脚本到可重用的转换组件,(2)开发一种统一的,基于循环/数组轴的方法来表达芯片上的通信,同时减少冗余在Loopy?S程序表示和转换,(3)探索高性能语言的设计空间,在执行放置和数据放置之间建立密切联系,(4)开发交互式程序转换和性能分析工具,并发现高性能计算劳动力培训的潜在影响,(5)演示所有开发的组件可以以实用和连贯的方式应用在一起。通过对研究生和本科生的教学,以及对该项目支持的学生和博士后的指导,PI为扩大人才库做出了贡献。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
A key capability in technical computing is the processing of large, regularly-shaped arrays of numbers by a wide variety of different processes. This facility is foundational in, for example, weather prediction, artificial intelligence, and image processing. Correspondingly, modern computing hardware has evolved advanced capabilities for carrying out such computations with high efficiency. Unfortunately, the process of adapting a desired process to a given piece of hardware thus far is costly, laborious, and error-prone. Differences of a factor of 50 in performance between a naive realization and a careful one is the rule, rather than the exception. Loopy, the subject of this project, attacks this problem by using human-guided, automated program rewriting. Loopy has demonstrated application impact in a number of applications ranging from the simulation of natural and engineering phenomena to neuroscience, where it has helped its users achieve higher performance with less effort. The present proposal concerns several important improvements that will contribute to making Loopy more effective and easier to apply, through enlarging the class of programs that Loopy can transform, improving the means by which Loopy represents on-chip communication, and permitting it to realize important basic operations that routinely pose difficulty in efficient implementation. An important component of the effort is making Loopy itself easy to use for its user community, through the realization of an interactive user interface, so that program transformations can be applied with the click of a mouse, rather than by writing computer code. The proposed advances will be demonstrated through a sample workload that is emblematic of many of the computational and software challenges faced in technical computing today.Multidimensional arrays (sometimes called 'tensors') are a foundational data structure for much of scientific computing, with applications ranging from weather prediction to deep learning, to image processing and computational neuroscience. Even the efficient execution of one of the simplest operations on arrays, matrix-matrix multiplication, poses considerable technical challenges on modern computers. Through a polyhedrally-based program transformation tool, the proposed software will provide separation between mathematical intent and the technical challenges of program optimization, allowing each task to be performed by a domain expert. In the proposed project, the PI will develop means for more efficient on-chip communication, code generation for prefix sums, reuse and abstraction in program transformation, increasing the ease of use in transformation discovery and performance analysis, and for expressing array computations in user programs. The PI will validate the proposed techniques through a challenging application with broad applicability. The intellectual merit of the proposed research lies in (1) mapping out and extending the landscape of transformation-based programming from one-off scripts to reusable transform components, (2) the development of a unifying, loop/array-axis-based approach to expressing on-chip communication while reducing redundancy in Loopy?s program representation and transformation, (3) exploring the design space of high-performance languages that establish a close link between execution placement and data placement, (4) the development of an interactive program transform and performance analysis tool, along with the discovery of potential implications for workforce training in high-performance computing, (5) a demonstration that all the developed components can be applied together in a practical and coherent manner. Through graduate and undergraduate teaching as well as mentoring of the students and postdocs supported by this project, the PI contributes to enlarging the talent pool.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Integral equation methods for the Morse-Ingard equations
Morse-Ingard 方程的积分方程方法
DOI: 10.1016/j.jcp.2023.112416
发表时间: 2023
期刊: Journal of Computational Physics
影响因子: 4.1
作者: [Wei, Xiaoyu, Klöckner, Andreas, Kirby, Robert C.]
通讯作者: Kirby, Robert C.
SHF: Small: Collaborative Research: Transform-to-Perform: Languages, Algorithms, and Solvers for Nonlocal Operators
CAREER: Towards General-Purpose, High-Order Integral Equation Methods for Computer Simulation in Engineering: Analysis, Algorithm Design, and Applications
Small: Collaborative Research: Transform-to-Perform: Languages, Algorithms, and Code Transformations for High-Performance FEM
Collaborative Research: Efficient High-Order Parallel Algorithms for Large-Scale Photonics Simulation
海外基金