Fluid: Resource-aware Hyperparameter Tuning Engine

Fluid: Resource-aware Hyperparameter Tuning Engine
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Peifeng Yu;Jiachen Liu;Mosharaf Chowdhury
Peifeng Yu;Jiachen Liu;Mosharaf Chowdhury
中科院分区:
其他
文献类型:
--
作者:
Peifeng Yu;Jiachen Liu;Mosharaf Chowdhury

文献摘要

相似文献

当前的超参数调优解决方案缺乏互补的执行引擎来有效利用分布式计算,从而忽略了 GPU 内和 GPU 间共享的可能性,从而导致资源利用率低下。在本文中,我们提出了 Fluid,一种通用的超参数调优执行引擎,可在超参数调优作业和集群资源之间进行协调。 Fluid 使用注水方法在此类作业中安排评估试验,以充分利用 GPU 内和 GPU 间粒度的资源,从而加快调整过程。通过将超参数调整作业抽象为 TrialGroup 序列,Fluid 可以提高各种超参数调整解决方案的性能。我们的实验表明,Fluid 可以将同步 BOHB 加速 100%,将 BOHB 和 ASHA 加速 30%,同时具有相似的最终精度。
Current hyperparameter tuning solutions lack complementary execution engines to efficiently leverage distributed computation, thus ignoring the possibility of intra-and inter-GPU sharing, which exhibits poor resource usage. In this paper, we present Fluid, a generalized hyperparameter tuning execution engine, that coordinates between hyperparameter tuning jobs and cluster resources. Fluid schedules evaluation trials in such jobs using a water-filling approach to make the best use of resources both at intra-and inter-GPU granularities to speed up the tuning process. By abstracting a hyperparameter tuning job as a sequence of TrialGroup, Fluid can boost the performance of diverse hyperparameter tuning solutions. Our experiments show that Fluid can speed up synchronous BOHB by 100% , and BOHB and ASHA by 30% while having similar final accuracy.