SHF: Small: Empirical Autotuning of Parallel Computation for Scalable Hybrid Systems
SHF: Small: Empirical Autotuning of Parallel Computation for Scalable Hybrid Systems
批准号:
1527706
负责人:
Jack Dongarra
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-07-15 至 2019-06-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Today, scientific and engineering computing is synonymous with parallel computing, and applications such as climate modeling, drug design, aircraft design, etc. utilize very large supercomputer installations, with power consumption measured in MegaWatts, and the cost of electricity measured in millions of dollars. At the same time, every parallel application requires some level of tuning to ensure that the software is mapped appropriately to the hardware. Otherwise, suboptimal performance can lead to lost cycles, kilowatt-hours, and, ultimately, dollars. Tuning the application by making repeated runs is also a wasteful option at very large scale. The DARE project addresses this problem by tuning the application through modeling and simulation of its behavior at very large scale, rather than actually running it. Therefore, resources required for tuning are marginal compared to those consumed in production runs. DARE is based on the observation that the same approach that replaces a wind tunnel with a computer simulation of the airfoil can be applied to the software itself. Two aspects of today's high-end computing landscape make the DARE work unique: 1) the prevalence of hardware accelerators, such as Graphics Processing Units and Xeon Phi co-processors, and 2) adoption of task-based, dynamic, work scheduling systems as an alternative to traditional, lock-step parallel programming models. In particular, DARE combines three components into a refinement loop: a hardware analysis component, a kernel modeling component, and a workload simulation component. The role of the hardware analysis component is to extract the basic hardware information, such as processing power and data link speed. The role of the kernel modeling component is to provide performance models of the serial kernels that constitute the building blocks of the parallel program. Finally, the role of the simulation component is to simulate large-scale parallel workloads.The hardware analysis component gathers the basic knowledge about the system, such as: the number of CPU sockets per shared memory node, the number of CPU cores in each socket, the cache hierarchy, existence of hyper-threading, number of NUMA nodes and proximity of CPUs to NUMA nodes, number of GPU accelerators or Xeon Phi co-processors and capacities of their device memories, and the topology and bandwidth of data links, both within each node (busses), and between nodes (network switches). Part of this knowledge can be gathered by using appropriate query APIs, such as hwloc, netloc, PAPI, and those provided in the CUDA SDK, OpenCL SDK, and Xeon Phi SDK. Synthetic tests can be used for parameters that cannot be established in this manner.Kernels are essentially the serial building blocks of parallel problems. Although kernels are usually characterized by serial control flow, most of the time they already rely on a high degree of data parallelism. Today's CPUs get most of their performance from SIMD parallelism, and GPUs get their performance from massive SIMT parallelism. The role of the kernel modeling component is two-fold: 1) to tune kernels for maximum performance at a given granularity, 2) to provide the kernel performance model as a function of granularity, which is changing to accommodate parallel execution.DARE turns to a stochastic time-stepping simulation in order to predict the performance of a dynamic runtime scheduler for two fundamental reasons: 1) Building good performance models on the basis of benchmarking actual parallel runs requires a significant number of runs with significant problem sizes, which is simply too time consuming. And 2), the impact of many tuning parameters is too complex to be modeled by sparsely sampling the tuning space and fitting simple curves / surfaces to the sample points. The answer to the problem is to replace the run with a time stepping simulation, where a given task-based scheduler is used for assigning tasks to cores, but instead of invoking actual kernel tasks, control is passed to a progress tracking simulation system, which relies on kernel performance models to simulate the execution of the tasks and produce a virtual trace of the simulated execution. The performance advantage is twofold: 1) Simulating a single run is much faster than actually making that run, and 2) Many simulations can be run in parallel allowing for fast sweeps through a large parameter search space.DARE replaces the standard waterfall autotuning process with a process that is incremental and iterative in nature. The power of the DARE approach lies in the mutual refinement loop, where each of the three phases is capable of massively pruning the search space for the other two. As a result, very high quality models can be built for a particular workload, since time is being spent refining the model for the conditions that actually apply, rather than sampling the search space in areas never touched at runtime.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Travel: Workshop on Clusters, Clouds, and Data Analytics for Scientific Computing 2024
-
批准号:2336813
-
项目类别:Standard Grant
-
资助金额:$2.5万
-
财政年份:2023
-
负责人:Jack Dongarra
-
依托单位:
Workshop on Clusters, Clouds, and Data Analytics for Scientific Computing
-
批准号:2001329
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:2020
-
负责人:Jack Dongarra
-
依托单位:
Workshop on Clusters, Clouds, and Data Analytics in Scientific Computing
-
批准号:1800946
-
项目类别:Standard Grant
-
资助金额:$1.93万
-
财政年份:2018
-
负责人:Jack Dongarra
-
依托单位:
Toward a common digital continuum platform for big data and extreme-scale computing (BDEC2)
-
批准号:1849625
-
项目类别:Standard Grant
-
资助金额:$20.34万
-
财政年份:2018
-
负责人:Jack Dongarra
-
依托单位:
Collaborative Research: ACI-CDS&E: Highly Parallel Algorithms and Architectures for Convex Optimization for Realtime Embedded Systems (CORES)
-
批准号:1709069
-
项目类别:Standard Grant
-
资助金额:$41.21万
-
财政年份:2017
-
负责人:Jack Dongarra
-
依托单位:
Workshop on Clusters, Clouds and Data Analytics in Scientific Computing
-
批准号:1606551
-
项目类别:Standard Grant
-
资助金额:$2.41万
-
财政年份:2016
-
负责人:Jack Dongarra
-
依托单位:
Collaborative Research: EMBRACE: Evolvable Methods for Benchmarking Realism through Application and Community Engagement
-
批准号:1535025
-
项目类别:Standard Grant
-
资助金额:$12.5万
-
财政年份:2015
-
负责人:Jack Dongarra
-
依托单位:
SI2-SSI: Collaborative Proposal: Performance Application Programming Interface for Extreme-Scale Environments (PAPI-EX)
-
批准号:1450429
-
项目类别:Standard Grant
-
资助金额:$212.64万
-
财政年份:2015
-
负责人:Jack Dongarra
-
依托单位:
CSR:Medium:Collaborative Research: SparseKaffe: high-performance, auto-tuned, energy-aware algorithms for sparse direct methods on modern heterogeneous architectures
-
批准号:1514286
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2015
-
负责人:Jack Dongarra
-
依托单位:
EAGER: Collaborative Research: Memristive Accelerator for Extreme Scale Linear Solvers
-
批准号:1548093
-
项目类别:Standard Grant
-
资助金额:$3.13万
-
财政年份:2015
-
负责人:Jack Dongarra
-
依托单位:
XPS: FULL: DSD: Collaborative Research: Rapid Prototyping HPC Environment for Deep Learning
-
批准号:1439052
-
项目类别:Standard Grant
-
资助金额:$38.25万
-
财政年份:2014
-
负责人:Jack Dongarra
-
依托单位:
SI2-SSI: Collaborative Research: Sustained Innovation for Linear Algebra Software (SILAS)
-
批准号:1339822
-
项目类别:Continuing Grant
-
资助金额:$119.0万
-
财政年份:2013
-
负责人:Jack Dongarra
-
依托单位:
SHF: Small: Bench-testing Environment for Automated Software Tuning (BEAST)
-
批准号:1320603
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2013
-
负责人:Jack Dongarra
-
依托单位:
Workshop on Clusters, Clouds and Grids for Scientific Computing
-
批准号:1226146
-
项目类别:Standard Grant
-
资助金额:$3.76万
-
财政年份:2012
-
负责人:Jack Dongarra
-
依托单位:
EAGER: PaRSEC: Parallel Runtime Scheduling and Execution Control
-
批准号:1244905
-
项目类别:Standard Grant
-
资助金额:$19.7万
-
财政年份:2012
-
负责人:Jack Dongarra
-
依托单位:
SHF: Small: Parallel Unified Linear algebra with Systolic ARrays (PULSAR)
-
批准号:1117062
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2011
-
负责人:Jack Dongarra
-
依托单位:
Supporting and Enhancing the HPC Challenge Benchmark for Hybrid-Multicore Computers
-
批准号:1038814
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2011
-
负责人:Jack Dongarra
-
依托单位:
Extending the Work of the International Exascale Software Project
-
批准号:1136509
-
项目类别:Standard Grant
-
资助金额:$9.98万
-
财政年份:2011
-
负责人:Jack Dongarra
-
依托单位:
Proposed Meeting Series: The Message Passing Interface Forum
-
批准号:1144042
-
项目类别:Standard Grant
-
资助金额:$7.03万
-
财政年份:2011
-
负责人:Jack Dongarra
-
依托单位:
Workshop on Clusters, Clouds and Grids for Scientific Computing
-
批准号:1032220
-
项目类别:Standard Grant
-
资助金额:$4.0万
-
财政年份:2010
-
负责人:Jack Dongarra
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: