ForkTail: a black-box fork-join tail latency prediction model for user-facing datacenter workloads
ForkTail: a black-box fork-join tail latency prediction model for user-facing datacenter workloads
复制标题
ForkTail:用于面向用户的数据中心工作负载的黑盒分叉连接尾部延迟预测模型
DOI:
10.1145/3208040.3208058
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Jiang, Hong
中科院分区:
文献类型:
--
作者:
Nguyen, Minh;Alesawi, Sami;Li, Ning;Che, Hao;Jiang, Hong
The workflows of the predominant user-facing datacenter services, including web searching and social networking, are underlaid by various Fork-Join structures. Due to the lack of understanding the performance of Fork-Join structures in general, today's datacenters often resort to resource overprovisioning, operating under low resource utilization, to meet stringent tail-latency service level objectives (SLOs) for such services. Hence, to achieve high resource utilization, while meeting stringent tail-latency SLOs, it is of paramount importance to be able to accurately predict the tail latency for a broad range of Fork-Join structures of practical interests.In this paper, we propose ForkTail, a black-box Fork-Join tail latency prediction model that covers a wide range of Fork-Join structures. In ForkTail, all Fork nodes are treated as black boxes, admitting both homogeneous and inhomogeneous cases, and different requests in the request flow are allowed to spawn different numbers of tasks forked to different numbers of Fork nodes. On the basis of the central limit theorem for queuing models under heavy load, we are able to arrive at a highly computational effective, empirical expression for the tail latency as a function of the means and variances of the task response times. Since this expression can be applied to request sub-flows at any granularities, it can be used for tail-latency prediction for services in a consolidated environment, where different services and applications may share the same datacenter cluster resources. Our extensive testing results based on model-based and trace-driven simulations, as well as a real-world case study in a cloud environment demonstrate that the expression can consistently predict the tail latency within 20% and 15% prediction errors at 80% and 90% load levels, respectively. Moreover, our sensitivity analysis demonstrates that such errors can be well compensated for with no more than 5% and 3% resource overprovisioning at these two load levels, respectively. This, together with its extremely low computational complexity, makes ForkTail a viable tool for both offline and online job scheduling and resource provisioning for user-facing datacenter applications.
登录
查看更多内容
DOI:
10.1109/71.946659
发表时间:
2001
期刊:
IEEE Trans. Parallel Distributed Syst.
影响因子:
--
作者:
R. Chen
通讯作者:
R. Chen
影响因子:
1
作者:
J. Köllerström
通讯作者:
J. Köllerström
DOI:
10.1145/2796314.2745859
发表时间:
2015-06
期刊:
ACM SIGMETRICS Performance Evaluation Review
影响因子:
--
作者:
Amr Rizk;Felix Poloczek;F. Ciucu
通讯作者:
Amr Rizk;Felix Poloczek;F. Ciucu
DOI:
10.1109/71.730531
发表时间:
1998
期刊:
IEEE Trans. Parallel Distributed Syst.
影响因子:
--
作者:
S. Balsamo;L. Donatiello;N. Dijk
通讯作者:
N. Dijk
DOI:
--
发表时间:
2015
期刊:
2015 15th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing
影响因子:
--
作者:
Z. Qiu;Juan F. Pérez
通讯作者:
Juan F. Pérez