Architecture of an On-Time Data Transfer Framework in Cooperation with Scheduler System

Architecture of an On-Time Data Transfer Framework in Cooperation with Scheduler System
复制标题

DOI:
10.1007/978-3-030-93571-9_13
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Kohei Yamamoto;Arata Endo;S. Date
Kohei Yamamoto;Arata Endo;S. Date
中科院分区:
其他
文献类型:
--
作者:
Kohei Yamamoto;Arata Endo;S. Date

文献摘要

相似文献

网络和物联网的技术进步为研究人员提供了新的方法或技术,可以使用物联网传感器和其他测量设备上观察到的最新数据进行数值分析和模拟。通常,大规模仿真需要高性能计算(HPC)系统。这种HPC系统在研究人员之间以共享的方式运行。因此,研究人员很难使用远程数据源生成的最新观测数据进行模拟。为了使研究人员能够利用新鲜的数据在远程数据源上进行计算,我们提出了一个及时的数据传输框架,使作业的执行与新鲜的数据在一个共享的HPC系统上的远程站点上生成的数据通过扩展SLURM调度。该框架包括两个功能:作业锁定和实时数据传输。该框架利用作业牵制功能,避免了调度算法对作业开始时间的重新安排。实时数据传输功能负责从远程站点到数据传输节点的数据传输。它试图在作业的固定开始时间完成数据传输。在本文中的评估表明,该框架可以保持数据的新鲜度高,最大限度地减少数据传输的作业等待时间。
Technological advancement in networking and IoT have given researchers new methods or techniques to perform numerical analysis and simulation with the latest data observed on IoT sensors and other measurement devices. In general, large-scale simulations necessitate high-performance computing (HPC) systems. Such HPC systems are operated in a shared manner among researchers. Therefore, it becomes inherently difficult for researchers to use the latest observation data generated on remote data sources for their simulations. To enable researchers to utilize fresh data on a remote data source for computation, we propose an on-time data transfer framework that enables the execution of jobs with fresh data generated on a remote site data on a shared HPC system by extending the SLURM scheduler. The proposed framework consists of two functions:Job pinningandOn-time data transfer. With the job pinning function, the proposed framework prevents the scheduling algorithm from rearranging the scheduled start time of jobs. The on-time data transfer function is in charge of data transfer from a remote site to the data transfer node. It attempts to complete the data transfer at just the time of the pinned start time of jobs. The evaluation in this paper indicates that the proposed framework can keep data freshness high and minimize the job waiting time for data transfer.