Towards Continually Learning Application Performance Models

Towards Continually Learning Application Performance Models
复制标题

DOI:
10.48550/arxiv.2310.16996
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Ray A. O. Sinurat
Ray A. O. Sinurat
中科院分区:
其他
文献类型:
--
作者:
Ray A. O. Sinurat

文献摘要

相似文献

基于机器学习的性能模型越来越多地用于构建关键的作业调度和应用优化决策。传统上,这些模型假设随着时间的推移收集更多的样本,数据分布不会改变。然而,由于生产HPC系统的复杂性和异构性,它们容易受到硬件降级、更换和/或软件补丁的影响,这可能导致数据分布的漂移,从而对性能模型产生不利影响。为此,我们开发了不断学习的性能模型,该模型考虑了分布漂移,减轻了灾难性遗忘,并提高了泛化能力。我们的最佳模型能够保持准确性,而不必学习系统变化造成的新数据分布,同时与朴素方法相比,整个数据序列的预测准确性提高了2倍。
Machine learning-based performance models are increasingly being used to build critical job scheduling and application optimization decisions. Traditionally, these models assume that data distribution does not change as more samples are collected over time. However, owing to the complexity and heterogeneity of production HPC systems, they are susceptible to hardware degradation, replacement, and/or software patches, which can lead to drift in the data distribution that can adversely affect the performance models. To this end, we develop continually learning performance models that account for the distribution drift, alleviate catastrophic forgetting, and improve generalizability. Our best model was able to retain accuracy, regardless of having to learn the new distribution of data inflicted by system changes, while demonstrating a 2x improvement in the prediction accuracy of the whole data sequence in comparison to the naive approach.