课题基金 / 基金详情

Decentralised, Large-scale Resource Management in Modern Data Centres

Decentralised, Large-scale Resource Management in Modern Data Centres
现代数据中心的分散式大规模资源管理
批准号:
EP/P009093/2
负责人:
Evangelia Kalyvianaki
金额:
$9.55万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

Evangelia Kalyvianaki的其他基金

相似基金

相关文献

中文摘要
翻译
现代全球信息技术(IT)和云基础设施的骨干由全球数据中心(DC)网络组成,每个数据中心都配备了数千台服务器。现代大型数据中心配备了50,000到100,000台服务器机器,并运行各种应用程序工作负载。报告显示,大约有300万个DC,包含1200万台服务器,运行着美国所有的在线业务。我们面临着一个DC环境,在服务器机器和应用程序的数量方面,应用程序部署的规模是前所未有的。现代发展中国家服务器的巨大规模极大地影响了发展中国家的资本和运营成本。资本成本包括发展中心设备(包括伺服器)的所有初期开支,而营运成本则用于发展中心的日常营运,包括电力消耗及管理人员的薪金。运行DC的成本是巨大的。报告显示,2015年全球在数据中心系统上的支出为1700亿美元,预计2016年将增长3%,达到1750亿美元。鉴于数据中心的高支出,现代数据中心以具有成本效益的方式运行至关重要,即服务器机器被运行的应用程序充分利用,应用程序被充分配置以满足其性能目标。然而,有许多报告显示,DC中的机器平均只有10-15%的CPU利用率。低利用率的主要原因是过度配置应用程序资源的做法,即使是最苛刻的应用程序工作负载需求,无论它们多么罕见。然而,由于工作量通常是时变的,变化未知,这种做法导致现代发展中国家资源利用率严重不足,从而导致发展中国家支出过剩。然而,从业者报告说,目前的管理框架不足以在云计算等大规模环境中执行可扩展的操作任务。因此,如何解决现代大规模DC中的资源管理问题,并在满足应用程序性能需求的同时提高整体资源利用率是一个开放的挑战。我们建议采用新的分散式资源管理方法,以解决区议会资源使用不足的问题。我们设想了一个分散的方案,其中资源调度器分布在DC上,每个调度器控制被称为集群的DC机器的子集的资源分配,即集群包含几百个服务器。使用集群调度器旨在及时提高集群内机器的有效利用率。所有数据中心服务器的全球资源规划是通过分散协调所有服务器来实现的。集群之间进行通信,以交换其集群的资源利用信息和应用程序性能信息,从而实现全局融合。为了提高整体利用率,目标是平衡所有集群的负载,同时避免热点和利用率不足。这项工作的新奇将是协调全球资源规划的一套分散的集群管理器。我们的目标是使用分布式优化和控制方法。这项工作的潜在影响是巨大的。我们预计在DC部门的经济和在人和知识领域的影响,因为拟议的工作将有助于IT管理员的技能的发展。最终受益者是社会,特别是云和IT应用程序的开发人员和最终用户。英国目前是欧洲最大的数据中心市场。拟议的研究有可能大大加强英国在重要的DC部门的地位,并影响其国际地位。
英文摘要
The backbone of modern, world-wide Information Technology (IT) and Cloud infrastructure consists of a global network of data centres (DCs) each equipped with thousands of server machines. Modern large DCs are equipped with 50,000 to 100,000 of server machines and run a diverse set of application workloads. Reports show that about three million DCs containing 12 million of server machines run all US online operations. We face a DC environment for application deployment of unprecedented scale with regards to the number of server machines and applications. The enormous scale of servers in modern DCs dramatically affects DCs' capital and operational costs. Capital costs include all initial spending for DC equipment, including server machines and operational costs are towards the DCs' daily operation including electricity consumption and personnel salaries for management. The costs for running DC are enormous. Reports show that the 2015 world-wide spending on DC systems was $170 billion and these are expected to grow by 3% for 2016 to $175 billion.Given the high DC expenditure it is of paramount importance that modern DCs operate in a cost-effective manner, i.e. server machines are fully utilised by running applications and applications are adequately provisioned to meet their performance goals. However, there are numerous reports showing that machines in DCs are on average only 10-15% CPU utilised. The main cause of low utilisation has been the practice of over-provisioning applications with resources to match even their most demanding application workload demands, however rare they might be. However, as workloads are typically time-varying with unknown variations, this practice has led to a dramatic under-utilisation of modern DC resources and consequently to an excess of DC expenditure. Futhermore, practitioners report that current management frameworks are inadequate to perform scalable operational tasks in large-scale environments such as the Cloud. It is therefore an open challenge how to tackle the resource management problem in modern large-scale DCs and increase the overall resource utilisation while satisfying applications' performance demands. We propose a new decentralised resource management approach to tackle the under-utilisation problem of DCs. We envisage a decentralised scheme where resource schedulers are distributed across the DC and each scheduler controls the resource allocation of a subset of the DC machines referred to as clusters, i.e. a cluster contains a few 100s of servers. The use of cluster schedulers aims to increase the effective utilisation of machines within a cluster in a timely fashion. Global resource planning across all DC servers is achieved through decentralised coordination of all schedulers. Schedulers communicate to exchange resource utilisation information of their clusters and application performance information for global convergence. To increase the overall utilisation, the goal is to balance the load across all clusters while avoiding hotspots and under-utiisation. The novelty of this work will be on the coordination of the distributed set of cluster schedulers for global resource planning. We aim to use a distributed optimisation and control approach. The potential impact of this work is huge. We anticipate an impact in the Economy of the DC sector and in the domains of People and Knowledge as the proposed work will assist the development of IT administrators' skills. The ultimate beneficiary is Society and in particular developers and end-users of Cloud and IT applications. UK currently holds the largest European data centre market. The proposed research has the potential to significantly strengthen the position of the UK in the important DC sector and impact its international position.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cdc45484.2021.9683763
发表时间: 2021-04
期刊: 2021 60th IEEE Conference on Decision and Control (CDC)
影响因子: --
作者: [Apostolos I. Rikos;Andreas Grammenos;Evangelia Kalyvianaki;C. Hadjicostis;Themistoklis Charalambous;K. Johansson]
通讯作者: Apostolos I. Rikos;Andreas Grammenos;Evangelia Kalyvianaki;C. Hadjicostis;Themistoklis Charalambous;K. Johansson
DOI: 10.1109/cdc45484.2021.9683229
发表时间: 2021-04
期刊: 2021 60th IEEE Conference on Decision and Control (CDC)
影响因子: --
作者: [Wei Jiang;Andreas Grammenos;Evangelia Kalyvianaki;Themistoklis Charalambous]
通讯作者: Wei Jiang;Andreas Grammenos;Evangelia Kalyvianaki;Themistoklis Charalambous
DOI: 10.1109/tnse.2023.3236214
发表时间: 2021-01
期刊: IEEE Transactions on Network Science and Engineering
影响因子: 6.6
作者: [Andreas Grammenos;Themistoklis Charalambous;Evangelia Kalyvianaki]
通讯作者: Andreas Grammenos;Themistoklis Charalambous;Evangelia Kalyvianaki
Decentralised, Large-scale Resource Management in Modern Data Centres
  • 批准号:
    EP/P009093/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $12.83万
  • 财政年份:
    2017
  • 负责人:
    Evangelia Kalyvianaki
  • 依托单位:
国内基金
海外基金
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    黄洛将
  • 依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    黄洛将
  • 依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
  • 批准号:
    12074246
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2020
  • 负责人:
    Yoshitomo Kamiya
  • 依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
  • 批准号:
    31972875
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    石江华
  • 依托单位: