Collaborative Research: SHF: Medium: Co-Optimizing Computation and Data Transformations for Sparse Tensors
Collaborative Research: SHF: Medium: Co-Optimizing Computation and Data Transformations for Sparse Tensors
批准号:
2106621
负责人:
David Lowenthal
金额:
$39.74万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-01-01 至 2026-12-31
中文摘要
稀疏张量计算是计算机辅助药物设计、欺诈检测和国家安全等重要应用的核心。及时执行这些应用程序可以提高用户的工作效率,并减少与每次执行相关的能源消耗。稀疏计算的特点是具有许多或大多数值为零的输入。为了避免存储和计算零值数据的效率低下,应用程序只存储非零数据,并使用辅助数据结构来恢复其位置。因此,稀疏张量计算表现出不可预测的内存访问模式,包括通过辅助数据结构的间接访问。因此,在今天的计算机体系结构上,稀疏张量计算的性能完全由数据的移动、通过内存系统和跨节点的移动所控制。数据移动在执行时间和能量消耗方面都是昂贵的。优化稀疏张量计算的数据移动作为高性能体系结构已经变得越来越多样化——传统的并行体系结构、用作并行加速器的图形处理器和复杂的内存系统——给软件开发人员带来了性能和生产力方面的挑战,他们最终要为每个平台编写特定于底层体系结构的代码。所提出的方法同时优化了数据在内存中的组织方式,计算的结构如何以减少数据移动的方式访问数据,以及计算和数据移动如何充分利用硬件体系结构的特性。由于数据的非零结构在程序执行之前是未知的,因此该方法还在其决策中检查运行时信息。由此产生的协同优化策略支持一种内聚方法,用于迭代地为广泛的稀疏计算制定调度和数据表示转换决策,并结合运行时适应性。这个项目正在开发一个编程框架,允许对稀疏计算进行高级规范,并对其进行优化以减少数据移动。它包括数据表示、数据布局和存储映射,以及用于稀疏计算的并行调度。它使用数据依赖关系、运行时信息和体系结构特性来完全绑定最终生成的代码。这种方法的目的是使稀疏张量计算能够处理依赖关系,如稀疏三角解和许多其他线性方程组的求解器,在稀疏张量上应用重排序,如Morton排序,以及稀疏张量数据表示的延迟绑定。研究的新颖和最有意义的方面包括:(1)可组合调度和数据转换,包括数据布局转换和存储映射;(2)检查器合成,用于数据表示、布局和存储映射之间的运行时数据转换,这些转换由外部函数组成;(3)支持数据相关张量计算;(4)部署在MLIR/LLVM编译器中的框架抽象。研究人员坚定地致力于扩大对计算机的参与,并制定了全面的计划,以吸引代表性不足的群体。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Sparse tensor computations are central to important applications including computer-assisted drug design, fraud detection, and national security. Timely execution of these applications improves user productivity and reduces the energy consumption associated with each execution. Sparse computations are characterized as having inputs where many or most values are zero. To avoid the inefficiency of storing and computing on zero-valued data, applications only store the nonzeros, with auxiliary data structures to recover their locations. As a result, sparse tensor computations exhibit unpredictable memory-access patterns that include indirection through the auxiliary data structures. Consequently, on today’s computer architectures, performance of sparse tensor computations is completely dominated by the movement of data, through the memory system and across nodes. Data movement is expensive both in terms of execution time and energy expenditure. Optimizing data movement of sparse tensor computations as high-performance architectures have become increasingly diverse — conventional parallel architectures, graphics processors used as parallel accelerators and complex memory systems — creates a performance and productivity challenge for software developers who end up writing low-level architecture-specific code for each platform. The proposed approach simultaneously optimizes how data is organized in memory, how the computation is structured to access the data in a way that reduces data movement, and how the computation and data movement make best use of features of the hardware architectures. Since the nonzero structure of the data is unknown until program execution, the approach also examines runtime information in its decisions. The resulting co-optimization strategy enables a cohesive approach for iteratively making scheduling and data representation transformation decisions for a wide range of sparse computations and incorporating runtime adaptations.This project is developing a programming framework that permits high-level specification of a sparse computation and optimizes it to reduce data movement. It composes data representations, data layouts and storage mappings, and parallel schedules for sparse computations. It employs data dependencies, runtime information, and architecture features to fully bind the final generated code. This approach is intended to enable handling sparse tensor computations with dependences such as sparse triangular solve and many other solvers for systems of linear equations, applying reorderings such as Morton ordering on sparse tensors, and late binding of sparse tensor data representations. The novel and most significant aspects of the research include: (1) composable schedule and data transformations, including data layout transformations and storage mapping; (2) inspector synthesis for runtime data transformations between data representations, layouts, and storage mappings, which are composed with external functions; (3) support for data-dependent tensor computations; and, (4) framework abstractions deployed in the MLIR/LLVM compiler.The researchers are strongly committed to broadening participation in computing and have comprehensive plans to engage the underrepresented groups.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Code Synthesis for Sparse Tensor Format Conversion and Optimization
稀疏张量格式转换和优化的代码综合
DOI:
--
发表时间:
2023
期刊:
International Symposium on Code Generation and Optimization
影响因子:
--
作者:
[Popoola, Tobi, Zhao, Tuowen, St. George, Aaron, Bhetwal, Kalyan, Strout, Michelle, Hall, Mary, Olschanowsky, Catherine]
通讯作者:
Olschanowsky, Catherine
DOI:
10.1145/3566054
发表时间:
2022-08
期刊:
ACM Transactions on Architecture and Code Optimization
影响因子:
1.6
作者:
[Tuowen Zhao;Tobi Popoola;Mary W. Hall;C. Olschanowsky;M. Strout]
通讯作者:
Tuowen Zhao;Tobi Popoola;Mary W. Hall;C. Olschanowsky;M. Strout
Runtime Composition of Iterations for Fusing Loop-carried Sparse Dependence
用于融合循环携带稀疏依赖的迭代的运行时组合
DOI:
10.1145/3581784.3607097
发表时间:
2023
期刊:
ACM
影响因子:
--
作者:
[Cheshmi, Kazem, Strout, Michelle, Mehri Dehnavi, Maryam]
通讯作者:
Mehri Dehnavi, Maryam
Collaborative Research: OAC Core: Improving Utilization of High-Performance Computing Systems via Intelligent Co-scheduling
-
批准号:2103511
-
项目类别:Standard Grant
-
资助金额:$25.03万
-
财政年份:2021
-
负责人:David Lowenthal
-
依托单位:
CSR: Rethinking System Software for Overprovisioned, High-Performance Computing Systems
-
批准号:1526015
-
项目类别:Standard Grant
-
资助金额:$49.0万
-
财政年份:2015
-
负责人:David Lowenthal
-
依托单位:
CSR: Small:Conductor: A Run-Time System for Exascale Computing
-
批准号:1216829
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2012
-
负责人:David Lowenthal
-
依托单位:
CSR-PSCE, SM: MPI-PPA: Improving Efficiency of Large-Scale Clusters Through Statistical Performance Prediction
-
批准号:0936251
-
项目类别:Continuing Grant
-
资助金额:$30.5万
-
财政年份:2009
-
负责人:David Lowenthal
-
依托单位:
CSR-PSCE, SM: MPI-PPA: Improving Efficiency of Large-Scale Clusters Through Statistical Performance Prediction
-
批准号:0834356
-
项目类别:Continuing Grant
-
资助金额:$32.0万
-
财政年份:2008
-
负责人:David Lowenthal
-
依托单位:
Collaborative Research: Efficient Detection and Alleviation of Scalability Problems
-
批准号:0429285
-
项目类别:Standard Grant
-
资助金额:$16.42万
-
财政年份:2004
-
负责人:David Lowenthal
-
依托单位:
SOFTWARE: Heterogeneous Cluster MPI: A System for Out-Of-Core, Heterogeneous Data Distribution
-
批准号:0234285
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:David Lowenthal
-
依托单位:
Instrumentation Grant for Research in Parallel and Distributed Computing
-
批准号:9986032
-
项目类别:Standard Grant
-
资助金额:$7.63万
-
财政年份:2000
-
负责人:David Lowenthal
-
依托单位:
Career: An Integrated Compiler/Run-Time System for Global Data Distribution
-
批准号:9733063
-
项目类别:Continuing Grant
-
资助金额:$20.01万
-
财政年份:1998
-
负责人:David Lowenthal
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: