Rearchitecting Linux Storage Stack for µs Latency and High Throughput

Rearchitecting Linux Storage Stack for µs Latency and High Throughput
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
Proceedings of the Sixteenth European Conference on Computer Systems
影响因子:
--
通讯作者:
Jaehyun Hwang;Midhul Vuppalapati;Simon Peter;R. Agarwal
Jaehyun Hwang;Midhul Vuppalapati;Simon Peter;R. Agarwal
中科院分区:
其他
文献类型:
--
作者:
Jaehyun Hwang;Midhul Vuppalapati;Simon Peter;R. Agarwal

文献摘要

相似文献

本文证明,即使数十个对延迟敏感的应用程序与以接近硬件容量的吞吐量执行读/写操作的吞吐量受限的应用程序争夺主机资源,使用Linux内核存储堆栈也可以实现µ s级延迟。此外,这种性能可以在不对应用程序、网络硬件、内核CPU处理器和/或内核网络堆栈进行任何修改的情况下实现。我们使用blk-switch(一种新的Linux内核存储堆栈架构)的设计、实现和评估来演示上述内容。blk-switch中的关键见解是,Linux的多队列存储设计,沿着多队列网络和存储硬件,使存储堆栈在概念上类似于网络交换机。blk-switch使用这种洞察力来调整计算机网络文献中的技术(例如,多个出口队列,单个请求的优先处理,负载平衡和交换机调度),以适应Linux内核存储堆栈。在各种场景下的blk-switch评估表明,它始终能够实现微秒级的平均延迟和尾部延迟(在第99和99 .第九章),同时允许应用程序近乎完美地利用硬件容量。
This paper demonstrates that it is possible to achieve µ s-scale latency using Linux kernel storage stack, even when tens of latency-sensitive applications compete for host resources with throughput-bound applications that perform read/write operations at throughput close to hardware capacity. Furthermore, such performance can be achieved without any modification in applications, network hardware, kernel CPU schedulers and/or kernel network stack. We demonstrate the above using design, implementation and evaluation of blk-switch , a new Linux kernel storage stack architecture. The key insight in blk-switch is that Linux’s multi-queue storage design, along with multi-queue network and storage hardware, makes the storage stack conceptually similar to a network switch. blk-switch uses this insight to adapt techniques from the computer networking literature ( e.g. , multiple egress queues, prioritized processing of individual requests, load balancing, and switch scheduling) to the Linux kernel storage stack. blk-switch evaluation over a variety of scenarios shows that it consistently achieves µ s-scale average and tail latency (at both 99 th and 99 . 9 th percentiles), while allowing applications to near-perfectly utilize the hardware capacity.