Deep Reinforcement Learning based Elasticity-compatible Heterogeneous Resource Management for Time-critical Computing

Deep Reinforcement Learning based Elasticity-compatible Heterogeneous Resource Management for Time-critical Computing
复制标题

DOI:
10.1145/3404397.3404475
复制
发表时间:
2020-08
期刊:
Proceedings of the 49th International Conference on Parallel Processing
影响因子:
--
通讯作者:
Zixia Liu;Liqiang Wang;Gang Quan
Zixia Liu;Liqiang Wang;Gang Quan
中科院分区:
其他
文献类型:
--
作者:
Zixia Liu;Liqiang Wang;Gang Quan

文献摘要

相似文献

快速生成的数据和大量的数据分析工作给底层计算设施带来了巨大的压力。因此,分布式多集群计算环境(诸如混合云)由于其在适应地理上分布的和潜在的基于云的计算资源方面的优势而提高了其必要性。形成这种环境的不同集群可能是异构的,也可能是资源弹性的。从分析的角度来看,随着对流应用和及时分析需求的不断增长,现在许多数据分析工作在时间紧迫性方面都是时间关键的。计算环境的总体工作负载可以是混合的,以包含时间关键型应用程序和一般应用程序。这些都需要一个有效的资源管理方法,能够理解计算环境和应用程序的功能。然而,系统的复杂性和高动态性极大地阻碍了传统的基于规则的方法的性能。在这项工作中,我们建议利用深度强化学习为异构分布式计算环境开发弹性兼容的资源管理,旨在减少错过时间截止日期的情况,同时保持较低的平均执行时间比。沿着强化学习,我们设计了一个深度模型,采用长短期记忆(LSTM)结构和部分模型共享的多目标学习机制。实验结果表明,该方法可以大大优于基线,并作为一个强大的资源管理的变化的工作负载。
Rapidly generated data and the amount magnitude of data analytical jobs pose great pressure to the underlying computing facilities. A distributed multi-cluster computing environment such as a hybrid cloud consequently raises its necessity due to its advantages in adapting geographically distributed and potentially cloud-based computing resources. Different clusters forming such an environment could be heterogeneous and may be resource-elastic as well. From analytical perspective, in accordance with increasing needs on streaming applications and timely analytical demands, many data analytical jobs nowadays are time-critical in terms of their temporal urgency. And the overall workload of the computing environment can be hybrid to contain both time-critical and general applications. These all call for an efficient resource management approach capable to apprehend both computing environment and application features. However, the added up complexity and high dynamics of the system greatly hinder the performance of traditional rule-based approaches. In this work, we propose to utilize deep reinforcement learning for developing elasticity-compatible resource management for a heterogeneous distributed computing environment, aiming for less occurrences of missing temporal deadline while maintaining low average execution time ratio. Along with reinforcement learning we design a deep model employing Long Short-Term Memory (LSTM) structure and partial model sharing for multi-target learning mechanism. The experimental results show that the proposed approach could greatly outperform baselines and serve as a robust resource management for variant workloads.