课题基金 / 基金详情

Decentralised, Large-scale Resource Management in Modern Data Centres

Decentralised, Large-scale Resource Management in Modern Data Centres
现代数据中心的分散式大规模资源管理
批准号:
EP/P009093/1
负责人:
Evangelia Kalyvianaki
金额:
$12.83万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

Evangelia Kalyvianaki的其他基金

相似基金

相关文献

中文摘要
翻译
现代全球信息技术(IT)和云基础设施的主干由全球数据中心(DC)网络组成,每个DC都配备了数千台服务器。现代大型数据中心配备了50,000到100,000台服务器,并运行各种应用程序工作负载。报告显示,大约300万个DC,包括1200万台服务器,运行着美国所有的在线操作。就服务器机器和应用程序的数量而言,我们面临的应用程序部署规模空前的DC环境。现代DC中服务器的巨大规模极大地影响了DC的资本和运营成本。资本成本包括DC设备的所有初始支出,包括服务器机器;运营成本用于DC的日常运营,包括电力消耗和管理人员工资。运营华盛顿的成本是巨大的。报告显示,2015年全球在DC系统上的支出为1,700亿美元,预计2016年将增长3%,达到1,750亿美元。鉴于DC的高额支出,现代DC以经济高效的方式运营至关重要,即服务器机器充分利用正在运行的应用程序,并且充分配置应用程序以实现其性能目标。然而,有大量报告显示,DC中的计算机平均只有10%-15%的CPU利用率。利用率低的主要原因是过度向应用程序提供资源,以满足其最苛刻的应用程序工作负载需求,无论这些需求可能多么罕见。然而,由于工作量通常是随时间变化的未知变化,这种做法导致现代数据中心资源的利用严重不足,从而导致数据中心支出过剩。此外,从业人员报告说,当前的管理框架不足以在云等大规模环境中执行可扩展的运营任务。因此,如何在满足应用程序性能需求的同时,解决现代大型分布式计算中心的资源管理问题,提高整体资源利用率,是一个有待解决的挑战。我们建议采用新的分散资源管理方法,以解决区议会使用率不足的问题。我们设想了一种分散的方案,其中资源调度器分布在DC上,并且每个调度器控制被称为集群的DC机器的子集的资源分配,即一个集群包含几个100个服务器。集群调度器的使用旨在及时提高集群内机器的有效利用率。所有DC服务器的全局资源规划是通过分散协调所有调度器来实现的。调度器通信以交换其集群的资源利用信息和用于全局收敛的应用性能信息。为了提高总体利用率,目标是平衡所有集群的负载,同时避免热点和利用率不足。这项工作的新颖性将是协调分布式集群调度器集以进行全局资源规划。我们的目标是使用分布式优化和控制方法。这项工作的潜在影响是巨大的。我们预计将对DC部门的经济以及人员和知识领域产生影响,因为拟议的工作将有助于发展IT管理员的技能。最终受益者是协会,尤其是云和IT应用程序的开发人员和最终用户。英国目前拥有最大的欧洲数据中心市场。拟议的研究有可能显著加强英国在重要的DC部门的地位,并影响其国际地位。
英文摘要
The backbone of modern, world-wide Information Technology (IT) and Cloud infrastructure consists of a global network of data centres (DCs) each equipped with thousands of server machines. Modern large DCs are equipped with 50,000 to 100,000 of server machines and run a diverse set of application workloads. Reports show that about three million DCs containing 12 million of server machines run all US online operations. We face a DC environment for application deployment of unprecedented scale with regards to the number of server machines and applications. The enormous scale of servers in modern DCs dramatically affects DCs' capital and operational costs. Capital costs include all initial spending for DC equipment, including server machines and operational costs are towards the DCs' daily operation including electricity consumption and personnel salaries for management. The costs for running DC are enormous. Reports show that the 2015 world-wide spending on DC systems was $170 billion and these are expected to grow by 3% for 2016 to $175 billion.Given the high DC expenditure it is of paramount importance that modern DCs operate in a cost-effective manner, i.e. server machines are fully utilised by running applications and applications are adequately provisioned to meet their performance goals. However, there are numerous reports showing that machines in DCs are on average only 10-15% CPU utilised. The main cause of low utilisation has been the practice of over-provisioning applications with resources to match even their most demanding application workload demands, however rare they might be. However, as workloads are typically time-varying with unknown variations, this practice has led to a dramatic under-utilisation of modern DC resources and consequently to an excess of DC expenditure. Futhermore, practitioners report that current management frameworks are inadequate to perform scalable operational tasks in large-scale environments such as the Cloud. It is therefore an open challenge how to tackle the resource management problem in modern large-scale DCs and increase the overall resource utilisation while satisfying applications' performance demands. We propose a new decentralised resource management approach to tackle the under-utilisation problem of DCs. We envisage a decentralised scheme where resource schedulers are distributed across the DC and each scheduler controls the resource allocation of a subset of the DC machines referred to as clusters, i.e. a cluster contains a few 100s of servers. The use of cluster schedulers aims to increase the effective utilisation of machines within a cluster in a timely fashion. Global resource planning across all DC servers is achieved through decentralised coordination of all schedulers. Schedulers communicate to exchange resource utilisation information of their clusters and application performance information for global convergence. To increase the overall utilisation, the goal is to balance the load across all clusters while avoiding hotspots and under-utiisation. The novelty of this work will be on the coordination of the distributed set of cluster schedulers for global resource planning. We aim to use a distributed optimisation and control approach. The potential impact of this work is huge. We anticipate an impact in the Economy of the DC sector and in the domains of People and Knowledge as the proposed work will assist the development of IT administrators' skills. The ultimate beneficiary is Society and in particular developers and end-users of Cloud and IT applications. UK currently holds the largest European data centre market. The proposed research has the potential to significantly strengthen the position of the UK in the important DC sector and impact its international position.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cdc45484.2021.9683763
发表时间: 2021-04
期刊: 2021 60th IEEE Conference on Decision and Control (CDC)
影响因子: --
作者: [Apostolos I. Rikos;Andreas Grammenos;Evangelia Kalyvianaki;C. Hadjicostis;Themistoklis Charalambous;K. Johansson]
通讯作者: Apostolos I. Rikos;Andreas Grammenos;Evangelia Kalyvianaki;C. Hadjicostis;Themistoklis Charalambous;K. Johansson
DOI: 10.1109/cdc45484.2021.9683229
发表时间: 2021-04
期刊: 2021 60th IEEE Conference on Decision and Control (CDC)
影响因子: --
作者: [Wei Jiang;Andreas Grammenos;Evangelia Kalyvianaki;Themistoklis Charalambous]
通讯作者: Wei Jiang;Andreas Grammenos;Evangelia Kalyvianaki;Themistoklis Charalambous
DOI: 10.1109/tnse.2023.3236214
发表时间: 2021-01
期刊: IEEE Transactions on Network Science and Engineering
影响因子: 6.6
作者: [Andreas Grammenos;Themistoklis Charalambous;Evangelia Kalyvianaki]
通讯作者: Andreas Grammenos;Themistoklis Charalambous;Evangelia Kalyvianaki
Decentralised, Large-scale Resource Management in Modern Data Centres
  • 批准号:
    EP/P009093/2
  • 项目类别:
    Research Grant
  • 资助金额:
    $9.55万
  • 财政年份:
    2017
  • 负责人:
    Evangelia Kalyvianaki
  • 依托单位:
国内基金
海外基金
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    黄洛将
  • 依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    黄洛将
  • 依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
  • 批准号:
    12074246
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2020
  • 负责人:
    Yoshitomo Kamiya
  • 依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
  • 批准号:
    31972875
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    石江华
  • 依托单位: