课题基金 / 基金详情

Collaborative Research: OAC Core: Enabling Extremely Fine-grained Parallelism on Modern Many-core Architectures

Collaborative Research: OAC Core: Enabling Extremely Fine-grained Parallelism on Modern Many-core Architectures
合作研究:OAC Core:在现代多核架构上实现极其细粒度的并行性
批准号:
2107548
负责人:
Ioan Raicu
金额:
$33.37万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-07-01 至 2024-06-30

项目摘要

项目成果

Ioan Raicu的其他基金

相似基金

相关文献

中文摘要
翻译
计算机系统正变得越来越复杂:具有多核处理器和通用图形处理器的多套接字系统有可能满足节点级苛刻应用程序的需求。可编程性和效率往往不容易兼顾,因为硬件在并行度上增长了几个数量级,一个芯片上有数千个计算单元。任务并行是一种重要的并行类型,它将计算分解为一组相互依赖的任务,这些任务可以在不同的计算单元上并发执行。为了实现强大的可伸缩性和高水平的有效并行性,当今的并行语言越来越需要支持过度分解(比内核多得多的任务),以提高性能,隐藏阻塞操作引起的延迟,并以其他方式实现最大的加速。通过在现代和未来硬件中不断增长的规模范围内实现对细粒度并行性的有效支持,可以预期并行程序员的生产力将得到提高。趋势表明,大多数Top500高性能计算系统可能会采用这项工作直接针对的硬件。该项目旨在开展一个具有广泛影响的分布式并行编程教育项目,鼓励学生在现实世界的挑战中实习,并为从研究到开源项目的技术转移铺平道路。特别强调妇女和代表性不足的少数民族的参与。这方面的教育将为科学家和工程师流畅地使用并行计算创造一个新的、更容易获得的基础。这项工作探索了新的数据结构和算法,允许在亚微秒时间尺度上实现细粒度并行的可扩展运行时和执行模型。pi在语言和运行时级别上的初步工作为实现这一目标提供了一条途径。该项目的目标是:1)统一运行时,支持以周期衡量的任务粒度:在不同节点硬件上设计、分析和实现高效细粒度计算的构建块;2)在一系列计算机体系结构的实际并行系统和应用内核的背景下评估这些构建块的性能;3)测量运行时对基准内核和实际应用程序的性能和可伸缩性影响;4)通过并行计算的新课程材料,将这项研究与从本科到研究生的教育项目相结合。这项高风险/高回报的研究旨在为各种规模的并行机器编程的便利性和效率带来变革性的改进。其贡献在于实现了高效的、隐式并行的高级语言,这些语言针对具有多核架构的单节点部署进行了优化,以支持以周期衡量的细粒度并行性,从而实现了一类全新的多任务计算应用程序。数据流体系结构使得隐式并行可以用编程模型处理,其影响可以与MATLAB、R和Python相媲美,另外的好处是相同的代码也可以在分布式系统或大规模HPC系统中运行。因此,科学家将能够编写一次程序,以任何合适的规模运行它,并让它无缝地为硬件的每个组件使用最合适的粒度。这项工作在数据流架构方面的创新将广泛适用于许多现有的并行编程系统,如OpenMP、Swift/Parsl和CUDA/OpenCL,在执行细粒度并行化的效率和在可能的情况下增加对隐式并行化的支持方面。目标硬件包括Intel/AMD x86、ThunderX/2 ARM、IBM Power9和NVIDIA/AMD gpu。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Computer systems are becoming increasingly complex: multisocket systems with many-core processors and general graphic processors have the potential to address the needs of demanding applications at the node level. Programmability and efficiency are often not easy to find together due to the hardware growing several orders of magnitude in degree of parallelism to thousands of computing units on a chip. Task parallelism is an important type of parallelism in which computation is broken down into a set of inter-dependent tasks which can be executed concurrently on various computing units. To achieve strong scaling and high levels of effective parallelism, there is a growing need in today's parallel languages with supporting over-decomposition (many more tasks than cores) in order to improve performance, hide latency caused by blocking operations, and otherwise achieve maximum speedup. By enabling the efficient support of fine-grained parallelism across the growing range of scales seen in modern and future hardware, it is expected that the productivity of parallel programmers will be enhanced. Trends show evidence that most of the Top500 high-performance computing systems will likely employ hardware that this work directly targets. The project aims to conduct a high-impact education program in distributed parallel programming with broad reach, encouraging student internships grounded in real-world challenges, and paving the way for technology transfer from research to open-source projects. Special emphasis is placed on engaging women and underrepresented minorities. This education facet will create a new and more accessible foundation for fluency in parallel computing for scientists and engineers.This work explores novel data-structures and algorithms that allow for scalable runtime and execution models for fine-grained parallelism at sub-microsecond timescales. Preliminary work by the PIs at the language and runtime levels suggests a path to achieving this. The project objectives are: 1) unifying runtime enabling task granularities measured in cycles: design, analysis, and implementation of building blocks for efficient fine-grained computing on diverse node hardware; 2) evaluating performance of these building blocks in the context of real parallel systems and application kernels on a range of computer architectures; 3) measuring performance and scalability impact of runtime on benchmark kernels and real applications; and 4) integrating this research with education programs from undergraduate to graduate levels through new course material on parallel computing. This high-risk/high-reward research is geared towards yielding transformative improvements in the ease and efficiency of programming parallel machines at every scale. The contributions lie in the realization of productive, implicitly parallel high-level languages optimized for single node deployments with many-core architectures to support fine-grained parallelism measured in cycles, enabling an entirely new class of many-task computing applications. The dataflow architecture makes implicit parallelism tractable with a programming model whose impact could rival that of MATLAB, R, and Python, with the added benefit that the same code could also run in a distributed system or large-scale HPC systems. Thus, the scientist would be able to write a program once, run it at any suitable scale, and have it seamlessly use the most appropriate granularity for each component of the hardware. This work’s innovations in dataflow architecture will be broadly applicable to a number of existing parallel programming systems such as OpenMP, Swift/Parsl, and CUDA/OpenCL, in terms of both efficiency in executing fine grained parallelism and adding support for implicit parallelism where possible. Target hardware includes Intel/AMD x86, ThunderX/2 ARM, IBM Power9, and NVIDIA/AMD GPUs.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: REU Site: BigDataX: From theory to practice in Big Data computing at eXtreme scales
  • 批准号:
    2150500
  • 项目类别:
    Standard Grant
  • 资助金额:
    $36.29万
  • 财政年份:
    2022
  • 负责人:
    Ioan Raicu
  • 依托单位:
REU Site: Collaborative Research: BigDataX: From theory to practice in Big Data computing at eXtreme scales
  • 批准号:
    1757964
  • 项目类别:
    Standard Grant
  • 资助金额:
    $32.31万
  • 财政年份:
    2018
  • 负责人:
    Ioan Raicu
  • 依托单位:
CRI: II-NEW: MYSTIC: Programmable Systems Research Testbed to Explore a Stack-WIde Adaptive System fabriC
  • 批准号:
    1730689
  • 项目类别:
    Standard Grant
  • 资助金额:
    $100.0万
  • 财政年份:
    2017
  • 负责人:
    Ioan Raicu
  • 依托单位:
REU Site: BigDataX: From Theory to Practice in Big Data Computing at Extreme Scales
  • 批准号:
    1461260
  • 项目类别:
    Standard Grant
  • 资助金额:
    $28.8万
  • 财政年份:
    2015
  • 负责人:
    Ioan Raicu
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)