NURD: Negative-Unlabeled Learning for Online Datacenter Straggler Prediction

NURD: Negative-Unlabeled Learning for Online Datacenter Straggler Prediction
复制标题

DOI:
10.48550/arxiv.2203.08339
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Yi Ding;Avinash Rao;Hyebin Song;R. Willett;Henry Hoffmann
Yi Ding;Avinash Rao;Hyebin Song;R. Willett;Henry Hoffmann
中科院分区:
其他
文献类型:
--
作者:
Yi Ding;Avinash Rao;Hyebin Song;R. Willett;Henry Hoffmann

文献摘要

被引文献

相似文献

数据中心执行大型计算作业,这些作业由较小的任务组成。当所有任务完成时,作业就完成了,因此散列任务(罕见但速度极慢的任务)是数据中心性能的主要障碍。准确预测离散井可以实现主动干预,使数据中心运营商能够在延迟作业之前减轻离散井的影响。虽然许多先前的工作应用机器学习来预测计算机系统的性能,但这些方法依赖于完整的标签——即所有可能行为的足够例子,包括离散和非离散——或者对潜在延迟分布的强假设——例如,是否为高斯分布。然而,在运行中的作业中,这些信息都是不可用的,直到掉队者已经延迟了作业时才显示出来。为了在没有标记的正例或延迟分布假设的情况下准确和早期地预测掉线者,本文提出了NURD,一种新的具有重加权和分布补偿的负-未标记学习方法,它只在负和未标记的流数据上进行训练。关键思想是利用非掉队者的已完成任务来训练一个预测器来预测未标记的运行任务的延迟,然后根据每个未标记任务的特征空间的加权函数对其预测进行重新加权。我们在谷歌和阿里巴巴的两条生产轨迹上对NURD进行了评估,发现与最佳基线方法相比,NURD在预测精度方面的F1得分提高了2- 11个百分点,在作业完成时间方面提高了2.0- 8.8个百分点。
Datacenters execute large computational jobs, which are composed of smaller tasks. A job completes when all its tasks finish, so stragglers -- rare, yet extremely slow tasks -- are a major impediment to datacenter performance. Accurately predicting stragglers would enable proactive intervention, allowing datacenter operators to mitigate stragglers before they delay a job. While much prior work applies machine learning to predict computer system performance, these approaches rely on complete labels -- i.e., sufficient examples of all possible behaviors, including straggling and non-straggling -- or strong assumptions about the underlying latency distributions -- e.g., whether Gaussian or not. Within a running job, however, none of this information is available until stragglers have revealed themselves when they have already delayed the job. To predict stragglers accurately and early without labeled positive examples or assumptions on latency distributions, this paper presents NURD, a novel Negative-Unlabeled learning approach with Reweighting and Distribution-compensation that only trains on negative and unlabeled streaming data. The key idea is to train a predictor using finished tasks of non-stragglers to predict latency for unlabeled running tasks, and then reweight each unlabeled task's prediction based on a weighting function of its feature space. We evaluate NURD on two production traces from Google and Alibaba, and find that compared to the best baseline approach, NURD produces 2--11 percentage point increases in the F1 score in terms of prediction accuracy, and 2.0--8.8 percentage point improvements in job completion time.