Minimizing draining waste through extending the lifetime of pilot jobs in Grid environments

Minimizing draining waste through extending the lifetime of pilot jobs in Grid environments
复制标题

通过延长网格环境中试点作业的生命周期,最大限度地减少排水浪费

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
F. Würthwein
F. Würthwein
中科院分区:
--
文献类型:
--
作者:
I. Sfiligoi;Thomas Martin;B. Bockelman;D. Bradley;F. Würthwein

文献摘要

被引文献

相似文献

计算领域正在加速向多核计算方向发展。如今,在单个物理节点上获得32个内核并不少见。因此,试点系统领域的压力越来越大,要求从纯粹的单核调度转变为允许多核作业。为了允许从单核用户作业逐步过渡到多核用户作业,设想试点作业将必须同时处理这两种用户作业,方法是一次从网格提供商请求多个核心,然后在运行时在用户作业之间划分它们。不幸的是,当前的网格生态系统只允许相对较短的试点作业的生命周期,需要频繁地耗尽,并且由于用户作业的生命周期不同而导致计算资源的相对浪费。因此,显著延长试点作业的生命周期是非常可取的,但必须不会对网格资源提供者造成任何不利影响。在本文中,我们提出了一种基于引导作业和网格提供者之间的通信的机制,该机制允许引导作业在有可用资源时运行较长的时间段,并且允许网格提供者在需要时在较短的时间内回收资源。我们还介绍了在几个美国的网格站点上使用上述机制运行原型系统的经验。
The computing landscape is moving at an accelerated pace to many-core computing. Nowadays, it is not unusual to get 32 cores on a single physical node. As a consequence, there is increased pressure in the pilot systems domain to move from purely single-core scheduling and allow multi-core jobs as well. In order to allow for a gradual transition from single-core to multi-core user jobs, it is envisioned that pilot jobs will have to handle both kinds of user jobs at the same time, by requesting several cores at a time from Grid providers and then partitioning them between the user jobs at runtime. Unfortunately, the current Grid ecosystem only allows for relatively short lifetime of pilot jobs, requiring frequent draining, with the relative waste of compute resources due to varying lifetimes of the user jobs. Significantly extending the lifetime of pilot jobs is thus highly desirable, but must come without any adverse effects for the Grid resource providers. In this paper we present a mechanism, based on communication between the pilot jobs and the Grid provider, that allows for pilot jobs to run for extended periods of time when there are available resources, but also allows the Grid provider to reclaim the resources in a short amount of time when needed. We also present the experience of running a prototype system using the above mechanism on a few US-based Grid sites.