课题基金 / 基金详情

Collaborative Research: OAC Core: Improving Utilization of High-Performance Computing Systems via Intelligent Co-scheduling

Collaborative Research: OAC Core: Improving Utilization of High-Performance Computing Systems via Intelligent Co-scheduling
合作研究:OAC Core:通过智能协同调度提高高性能计算系统的利用率
批准号:
2103511
负责人:
David Lowenthal
金额:
$25.03万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-01 至 2024-08-31

项目摘要

项目成果

David Lowenthal的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This project is aimed at increasing efficiency of high-performance computing systems by scheduling multiple jobs on the same set of nodes in a system, generally called co-scheduling. This is a break from current practice in which nodes are dedicated to one job at a time, which results in predictable execution time but inefficient use of system resources. To make this practical, the project will develop analyses to determine how to carry out co-scheduling such that overall system efficiency is improved while the performance impact on individual applications is minimized. In particular, the goal is to co-schedule jobs that can co-exist without contending for similar resources on the nodes. The work done in this project will help achieve better efficiency on high-performance systems, which will impact application domains such as climate/weather, renewable energy, and national security. The work will be implemented and validated on systems at Lawrence Livermore and Sandia National Laboratories and then transitioned into software that will be used at these national laboratories. The project will also have an impact on education by integrating the techniques in this research into courses covering parallel and distributed computing at the PIs' institutions. In addition, the project will take place at two Hispanic-serving institutions, and the PIs have a history of advising under-represented students; the project will broaden participation in computing by recruiting Hispanic undergraduates to work on the project and sending them to national laboratories for internships.The long-standing abstraction at high-end computing facilities is one of a submitted job being allocated a set of dedicated nodes. However, this makes systems much less efficient, as there are more per-node resources that will often be used inefficiently. In addition, the demand for high-end systems is increasing and dedicating nodes to jobs can increase job turnaround time and decrease overall system throughput. One way to address this problem is for supercomputer centers to break from the current common practice of assigning each job a private, isolated portion of a supercomputer. The intellectual merit of the project is three-fold. First, novel profile analyses will be developed that will reveal the effects on jobs due to sharing nodes. Second, novel statistical projection techniques will be developed that predict scaling behavior of jobs that are utilizing shared nodes. Third, new job-level scheduling techniques will be designed that use the interference analysis and projections to choose a set of shared nodes that will lead to good job turnaround time and maximize system throughput. The broader impact of the project is multifold. This project will help achieve better efficiency on high-performance systems, which will benefit a broad range of applications that includes climate/weather prediction, nuclear energy, and national security. Through a long-standing collaboration with both Lawrence Livermore and Sandia National Laboratories, the PIs will implement and validate the techniques on LLNL and SNL systems as well as transition the techniques into future resource managers at the national laboratories. In addition, both PIs will broaden participation in computing by recruiting Hispanic undergraduates to work on the project and sending them to national labs for internships.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2023
期刊: Workshop on Job Scheduling Strategies for Parallel Processing
影响因子: --
作者: [Hall, Jason, Lathi, Arjun, Lowenthal, David K, Patki, Tapasya]
通讯作者: Patki, Tapasya
Collaborative Research: SHF: Medium: Co-Optimizing Computation and Data Transformations for Sparse Tensors
  • 批准号:
    2106621
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $39.74万
  • 财政年份:
    2022
  • 负责人:
    David Lowenthal
  • 依托单位:
CSR: Rethinking System Software for Overprovisioned, High-Performance Computing Systems
  • 批准号:
    1526015
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.0万
  • 财政年份:
    2015
  • 负责人:
    David Lowenthal
  • 依托单位:
CSR: Small:Conductor: A Run-Time System for Exascale Computing
  • 批准号:
    1216829
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2012
  • 负责人:
    David Lowenthal
  • 依托单位:
CSR-PSCE, SM: MPI-PPA: Improving Efficiency of Large-Scale Clusters Through Statistical Performance Prediction
  • 批准号:
    0936251
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $30.5万
  • 财政年份:
    2009
  • 负责人:
    David Lowenthal
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)