Multicore job scheduling in the Worldwide LHC Computing Grid

Multicore job scheduling in the Worldwide LHC Computing Grid
复制标题

全球 LHC 计算网格中的多核作业调度

DOI:
10.1088/1742-6596/664/6/062016
复制
发表时间:
2015
期刊:
Conference Series
影响因子:
--
通讯作者:
Forti A
Forti A
中科院分区:
--
文献类型:
--
作者:
Forti A

文献摘要

参考文献

被引文献

相似文献

在LHC首次成功运行后,数据采集计划于2015年夏季重新启动,实验条件导致数据量和事件复杂性增加。为了处理这种情况下产生的数据,并利用当前CPU的多核架构,LHC实验已经开发了并行化的数据重建和模拟软件。然而,他们的计算工作的很大一部分仍然被期望作为单核任务执行。因此,不同的资源需求的工作将分布在全球大型强子对撞机计算网格(WLCG),使工作负载调度本身就是一个复杂的问题。为了应对这一挑战,WLCG多核部署工作组已经成立,以协调实验和WLCG站点的共同努力。其主要目标是确保不同的LHC虚拟组织(VO)的方法的收敛,以最大限度地利用共享资源,以满足他们的新的计算需求,最大限度地减少任何效率低下的调度机制,而不施加不必要的复杂性的网站管理他们的资源的方式。本文介绍了与上述主题相关的工作组的活动和进展,包括如何最好地使用不同的批处理系统技术的关键站点的经验,工作量提交工具的实验和不同的建议作业提交策略的规模测试中获得的知识的演变。
After the successful first run of the LHC, data taking is scheduled to restart in Summer 2015 with experimental conditions leading to increased data volumes and event complexity. In order to process the data generated in such scenario and exploit the multicore architectures of current CPUs, the LHC experiments have developed parallelized software for data reconstruction and simulation. However, a good fraction of their computing effort is still expected to be executed as single-core tasks. Therefore, jobs with diverse resources requirements will be distributed across the Worldwide LHC Computing Grid (WLCG), making workload scheduling a complex problem in itself. In response to this challenge, the WLCG Multicore Deployment Task Force has been created in order to coordinate the joint effort from experiments and WLCG sites. The main objective is to ensure the convergence of approaches from the different LHC Virtual Organizations (VOs) to make the best use of the shared resources in order to satisfy their new computing needs, minimizing any inefficiency originated from the scheduling mechanisms, and without imposing unnecessary complexities in the way sites manage their resources. This paper describes the activities and progress of the Task Force related to the aforementioned topics, including experiences from key sites on how to best use different batch system technologies, the evolution of workload submission tools by the experiments and the knowledge gained from scale tests of the different proposed job submission strategies.
ATLAS AthenaMP 的多核作业提交和网格资源调度
DOI: --
发表时间: 2012
期刊:
影响因子: --
作者:
David Crooks;P. Calafiura;R. Harrington;M. Jha;T. Maeno;S. Purdie;H. Severini;S. Skipsey;V. Tsulaia;R. Walker;A. Washbrook
通讯作者: A. Washbrook
CMS 工作负载管理向多核作业支持的演变
DOI: 10.1088/1742-6596/664/6/062046
发表时间: 2015
期刊: Journal of Physics: Conference Series
影响因子: --
作者:
A. P. Yzquierdo;J. Hernández;F. A. Khan;J. Letts;K. Majewski;A M Rodrigues;A. McCrea;E. Vaandering
通讯作者: E. Vaandering
在共享多用途集群上调度多核工作负载
DOI: --
发表时间: 2015
期刊:
影响因子: --
作者:
J. Templon;C. Acosta;J Flix Molina;A. Forti;A. P. Yzquierdo;R. Starink
通讯作者: R. Starink
LHC Run 2 的 CMS TierO 进入云和网格
DOI: 10.1088/1742-6596/664/3/032014
发表时间: 2015
期刊: Journal of Physics: Conference Series
影响因子: --
作者:
D. Hufnagel
通讯作者: D. Hufnagel