课题基金 / 基金详情

Collaborative Research: SHF:SMALL: Compile-Parallelize-Schedule-Retarget-Repeat (EASER) Paradigm for Dealing with Extreme Heterogeneity

Collaborative Research: SHF:SMALL: Compile-Parallelize-Schedule-Retarget-Repeat (EASER) Paradigm for Dealing with Extreme Heterogeneity
合作研究:SHF:SMALL:处理极端异构性的编译-并行化-调度-重定向-重复 (EASER) 范式
批准号:
2146852
负责人:
Gagan Agrawal
金额:
$25.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-06-15 至 2023-07-31

项目摘要

项目成果

Gagan Agrawal的其他基金

相似基金

相关文献

中文摘要
翻译
计算中的异构性是指在一个计算系统甚至集群的一个节点中存在各种各样的设备。许多技术趋势使得高性能计算(HPC)不可避免地出现高度异构,从而导致了多个方向的研究。传统的调度问题(指的是获取一组要执行的程序并将它们映射到可用资源)在这种异构性的存在下变得更加复杂,因为调度程序也需要与编译器交互。这个项目的目标是根据这些发展考虑应用程序执行的新范例,并在开发执行时间、编译、并行化和调度的预测方面进行研究。传统上,决定(可能是手动的)如何并行化应用程序、编译和集群级调度是顺序且独立地完成的。研究人员认为,当试图优化多租户异构集群时,他们的孤立处理是不可接受的。相反,研究人员设想的需求可以被称为EASER——compilE-pArallelize-Schedule-rEtarget-Repeat。为了详细说明这一愿景,在EASER范例中,编译器首先将核心功能映射到特定的设备,生成执行时间的预测,并将其输入到并行化方法选择模块,然后它们一起生成最终的可执行文件。随后,将此二进制文件提交给调度器,调度器评估作业队列,并可能建议替代配置/设备。如果是这样,将调用重定向模块,这可能会导致上述步骤的重复。该项目在执行新兴机器学习(ML)工作负载的集群上下文中开发、支持和评估EASER框架。1)编译器驱动的性能预测——它包括一种新的策略,该策略包括预测SIMD/VLIW性能的通用模型和基于操作符分类的开发内存层次性能模型的方法。2)集成作业调度和并行化策略选择——在性能预测模型的基础上,通过包括参数化和增量并行化策略选择方法并积极减少调度方法中的搜索空间,将这两个(传统上独立的)模块集成在一起。3)重定向编译器——通过将优化分类为依赖于架构的或独立的,将开发用于ML工作负载的重定向编译器。该项目还将对教育和人力资源发展作出若干贡献。两位研究者将介绍计算机系统和机器学习交叉的课程(材料),引起人们对计算机系统教育中与机器学习相关的工作量的关注。每所大学的大部分资金将用于支持博士生的研究,他们将接受跨传统(子)领域工作的培训。两位研究人员都坚定地致力于增加计算机领域的多样性,并在他们的研究项目中监督代表性不足的群体的成员。在两国大学现有联系的基础上,他们将进一步努力提高各个层面的多样性。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Heterogeneity in computing refers to having a variety of devices present within one computing system or even within one node of a cluster. A number of technological trends are making a high degree of heterogeneity inevitable in High Performance Computing (HPC), leading to research along many directions. The traditional scheduling problem, which refers to taking a set of programs to be executed and mapping them to the available resources, becomes more complicated in the presence of such heterogeneity, as the schedulers need to interact with the compiler also. The goal of this project is to consider new paradigms for application execution in view of these developments and conduct research in developing predictions of execution times, compilation, parallelization, and scheduling. Traditionally, deciding (likely manually) how an application is to be parallelized, compilation, and cluster-level scheduling are done sequentially and independently. The investigators posit that their isolated treatment is not going to be acceptable when one tries to optimize for multi-tenant heterogeneous clusters. Instead, the investigators envision a requirement that can be referred to as EASER -- compilE-pArallelize-Schedule-rEtarget-Repeat. To elaborate on the vision, in the EASER paradigm the compiler first maps the core functions to a specific device, generating predictions of execution time that are input to the parallelization approach selection module, and together they produce a final executable. Subsequently, this binary is presented to the scheduler, which assesses the job queue and might suggest alternative configuration(s)/device(s). If so, a retargeting module is to be invoked, leading to a potential repetition of the above steps. This project develops, supports, and evaluates the EASER framework in the context of a cluster that executes emerging machine learning (ML) workloads. Research is proposed in the following areas: 1) Compiler-Driven Performance Prediction -- It includes a novel strategy that comprises a general model for predicting SIMD/VLIW performance and an operator classification based approach to developing a memory hierarchy performance model. 2) Integrated Job Scheduling and Parallelization Strategy Selection -- Building on the performance prediction models, these two (conventionally independent) modules are integrated, by including parameterized and incremental parallelization strategy selection methods and aggressively reducing the search space in scheduling methods. 3) Retargeting Compiler -- By classifying optimizations as either architecture-dependent or independent, a retargeting compiler for ML workloads will be developed. This project will also make several contributions to education and human resource development. Both investigators will be introducing course(s) (material) at the intersection of computer systems and machine learning, bringing attention to ML-related workloads in computer systems education. A majority of funds at each University will be used to support Ph.D. students in their research, who will be trained to work across traditional (sub-) areas. Both investigators are strongly committed to increasing diversity in computing fields and have a strong record of supervising members of underrepresented groups in their research programs. Building on their Universities' existing connections, they will be further working on improving diversity at all levels.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3572848.3577486
发表时间: 2023-02
期刊: Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming
影响因子: --
作者: [Yang Xia;Peng Jiang;G. Agrawal;R. Ramnath]
通讯作者: Yang Xia;Peng Jiang;G. Agrawal;R. Ramnath
Collaborative Research: CNS Core: Small: A Compilation System for Mapping Deep Learning Models to Tensorized Instructions (DELITE)
Collaborative Research: CNS Core: Small: A Compilation System for Mapping Deep Learning Models to Tensorized Instructions (DELITE)
OAC Core: SHF: SMALL: ICURE -- In-situ Analytics with Compressed or Summary Representations for Extreme-Scale Architectures
SHF: Small: K-Way Speculation for Mapping Applications with Dependencies on Modern HPC Systems
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)