High Performance Computing in Science and Engineering - 4th International Conference, HPCSE 2019, Karolinka, Czech Republic, May 20-23, 2019, Revised Selected Papers

High Performance Computing in Science and Engineering - 4th International Conference, HPCSE 2019, Karolinka, Czech Republic, May 20-23, 2019, Revised Selected Papers
复制标题

科学与工程中的高性能计算 - 第四届国际会议,HPCSE 2019,捷克共和国卡罗林卡,2019 年 5 月 20-23 日,修订精选论文

DOI:
10.1007/978-3-030-67077-1_7
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Jaros M
Jaros M
中科院分区:
--
文献类型:
--
作者:
Jaros M

文献摘要

相似文献

在将复杂的生物医学工作流自动卸载到云和高性能设施的过程中,执行参数的估计占据了中心位置。由于普通用户对工作流中特定任务的性能特性没有或非常有限的知识,因此调度系统必须具有选择适当数量的计算资源的能力,例如,计算节点、GPU或处理器内核,并估计执行时间和成本。所提出的方法考虑了一组固定的可执行文件,这些可执行文件可用于创建自定义工作流,并收集成功计算的任务的性能数据。由于工作流可能在输入数据的结构和大小方面有所不同,因此只能通过搜索性能数据库并在类似任务之间进行插值来获得执行参数。本文表明,它是有可能预测的执行时间和成本具有较高的信心。如果在性能数据库中找到任务参数,则平均插值误差保持在2.29%以下。如果只找到类似的任务,平均插值误差可能会增长到15%。尽管如此,这仍然是一个可以接受的错误,因为集群性能也可能以百分比的顺序变化。
Estimation of execution parameters takes centre stage in automatic offloading of complex biomedical workflows to cloud and high performance facilities. Since ordinary users have no or very limited knowledge of the performance characteristics of particular tasks in the workflow, the scheduling system has to have the capabilities to select appropriate amount of compute resources, e.g., compute nodes, GPUs, or processor cores and estimate the execution time and cost.The presented approach considers a fixed set of executables that can be used to create custom workflows, and collects performance data of successfully computed tasks. Since the workflows may differ in the structure and size of the input data, the execution parameters can only be obtained by searching the performance database and interpolating between similar tasks. This paper shows it is possible to predict the execution time and cost with a high confidence. If the task parameters are found in the performance database, the mean interpolation error stays below 2.29%. If only similar tasks are found, the mean interpolation error may grow up to 15%. Nevertheless, this is still an acceptable error since the cluster performance may vary on order of percent as well.