Provisioning computational resources using virtual machines and leases

Provisioning computational resources using virtual machines and leases
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
Ian T Foster;Borja Sotomayor Basilio
Ian T Foster;Borja Sotomayor Basilio
中科院分区:
其他
文献类型:
--
作者:
Ian T Foster;Borja Sotomayor Basilio

文献摘要

被引文献

相似文献

多年来,对计算资源的需求已经成为科学和工业的基本要求。在许多情况下,这种需求是暂时的:用户可能只需要在明确定义的任务的持续时间内使用计算资源。例如,科学家可能需要大量计算机来运行模拟几个小时,但可能在任何特定时间都不需要这些计算机(只要它们在合理的时间内可用)。大学教师可能希望在课程的实验期间,在一周中的非常特定的时间,并使用特定的软件配置,为学生提供一组计算机。一家电信公司可能拥有托管多个网站的现有基础设施,但在网络流量意外增加期间,可能需要用额外的资源补充该基础设施,这意味着这些资源必须在几乎没有事先通知的情况下立即提供。这些短暂的资源使用场景提出了如何有效地提供共享计算资源的问题。这个问题已经研究了几十年,导致方法往往高度专业化,以特定的使用场景。例如,如何在共享集群上运行多个作业的问题已经得到了广泛的研究,导致了诸如TORQUE/MAUI、Sun网格引擎、LoadLeveler等作业管理系统系统,它们可以高效地对作业请求进行排队和优先排序(在这些系统中,效率是根据各种指标来定义的,包括等待时间和资源利用率)。这样的系统将满足希望在几个小时内运行模拟的科学家的要求,但另一方面,上述大学教师和电信公司将因工作管理系统和通常在工作管理中使用的效率度量而受到不利影响。相反,其他资源配置方法并不特别适合面向作业的计算。因此,没有通用的解决方案可以同时提供满足不同使用场景的要求的资源,例如上面提到的那些,以协调每个场景中的不同效率度量。更具体地说,我的大部分工作都是由尽力而为的资源要求和提前预订资源要求的组合所驱动的,其中用户需要计算资源,但愿意等待(可能设置了最后期限),提前预订资源要求必须在特定时间可用。在前者中,效率通常是根据等待时间(或类似的指标,如周转时间或减速)或吞吐量来衡量的,而后者通常关心的是在没有中断的情况下准确地在商定的时间内提供所请求的资源,两者都关心硬件资源的最大化使用以及可能的金钱利润。尽管已经分别研究了尽力而为和提前预留,但已知两者的组合会产生利用率问题,在实践中不鼓励这样做。在本论文中,我开发了一种能够同时高效地支持多种资源配置场景的资源配置模型和体系结构,主要集中在上面提到的尽力而为和提前预留的情况,并支持基于租约的模型,在该模型中,租约被实现为虚拟机(VM)。本文的主要贡献在于:1.提出了一种以租约为基础抽象、以虚拟机为实现载体的资源供应模型和体系结构。2.缓解在调度预留时通常遇到的利用率问题的租赁调度算法。3.一种用于使用虚拟机所涉及的各种开销的模型和算法,所述算法(A)即使在存在该开销的情况下也允许满足租赁条款,以及(B)在某些情况下减轻该开销。4.基于价格的租赁准入政策,表明自适应定价策略在某些情况下可以产生比其他基准定价策略更多的收入,但XIV通过使用更少的资源来实现这一点,从而使资源提供者有更多的过剩能力,有可能出售给其他用户。作为技术贡献,我还提供了Haizea(http://haizea.cs.uchicago.edu/),),这是本文所描述的体系结构和算法的开源参考实现。
The need for computational resources has, over the years, become a fundamental requirement in both science and industry. In many cases, this need is transient: a user may only require computational resources for the duration of a well-defined task. For example, a scientist could require a large number of computers to run a simulation for just a few hours, but might not need those computers at any specific time (as long as they are made available in a reasonable amount of time). A college instructor may want to make a cluster of computers available to students during the course's lab sessions, at very specific times during the week, and with a specific software configuration. A telecommunications company could possess an existing infrastructure that hosts a number of websites, but may need to supplement that infrastructure with additional resources during periods of unforeseen increased web traffic, meaning those resources have to made available right away with very little advance notice. These transient resource usage scenarios pose the problem of how to provision shared computational resources efficiently. This problem has been studied for decades, resulting in approaches that tend to be highly specialized to specific usage scenarios. For example, the problem of how to run multiple jobs on a shared cluster has been extensively studied, resulting in job management systems systems like Torque/Maui, Sun Grid Engine, LoadLeveler, and many others, that can queue and prioritize job requests efficiently (in these systems, efficiency is defined in terms of a variety of metrics, including waiting times and resource utilization). Such a system would meet the requirements of the scientist wanting to run simulations during a few hours but, on the other hand, the college instructor and the telecommunications company mentioned above would be ill-served by a job management system and the efficiency metrics typically used in job management. Conversely, other resource provisioning approaches are not particularly well suited for job-oriented computations. Thus, there is no general solution that can provision resources meeting the requirements of different usage scenarios simultaneously, such as those mentioned above, reconciling the different measures of efficiency in each scenario. More specifically, much of my work is motivated by the combination of best-effort resource requirements, where a user needs computational resources but is willing to wait for them (possibly setting a deadline), and advance reservation resource requirements, where the resources must be available at a specific time. In the former, efficiency is typically measured in terms of waiting times (or similar metrics such as turnaround times or slowdowns) or throughput, while the latter is usually concerned with providing the requested resources at exactly the agreed-upon times without interruption, and both are concerned with maximizing the use of hardware resources and possibly monetary profit. Although both best-effort and advance reservation provisioning have been studied separately, the combination of both is known to produce utilization problems and is discouraged in practice. In this dissertation I develop a resource provisioning model and architecture that can support multiple resource provisioning scenarios efficiently and simultaneously, with an initial focus on the best-effort and advance-reservation cases mentioned above, and arguing in favour of a lease-based model, where leases are implemented as virtual machines (VMs). The main contributions of this dissertation are: 1. A resource provisioning model and architecture that uses leases as a fundamental abstraction and virtual machines as an implementation vehicle. 2. Lease scheduling algorithms that mitigate the utilization problems typically encountered when scheduling advance reservations. 3. A model for the various overheads involved in using virtual machines, and algorithms that (a) allow leaseterms to be met even in the presence of this overhead, and (b) mitigate this overhead in some cases. 4. Price-based policies for lease admission, showing that an adaptive pricing strategy can, in some cases, generate more revenue than other baseline pricing strategies, but does xiv so by using fewer resources, thus giving resource providers more excess capacity that can potentially be sold to other users. As a technological contribution, I also present Haizea (http://haizea.cs.uchicago.edu/), an open source reference implementation of the architecture and algorithms described in this dissertation.