Towards low-latency I/O services for mixed workloads using ultra-low latency SSDs

Towards low-latency I/O services for mixed workloads using ultra-low latency SSDs
复制标题

DOI:
10.1145/3524059.3532378
复制
发表时间:
2022-06
期刊:
Proceedings of the 36th ACM International Conference on Supercomputing
影响因子:
--
通讯作者:
Mingzhe Liu;Haikun Liu;Chencheng Ye;Xiaofei Liao;Hai Jin;Yu Zhang;Ran Zheng;Liting Hu
Mingzhe Liu;Haikun Liu;Chencheng Ye;Xiaofei Liao;Hai Jin;Yu Zhang;Ran Zheng;Liting Hu
中科院分区:
其他
文献类型:
--
作者:
Mingzhe Liu;Haikun Liu;Chencheng Ye;Xiaofei Liao;Hai Jin;Yu Zhang;Ran Zheng;Liting Hu

文献摘要

被引文献

相似文献

低延迟I/O服务对于云数据中心的延迟敏感工作负载至关重要。尽管高级SSD(例如Intel Optane SSD)可以在设备层提供超低延迟,但通过I/O堆栈之间的各种工作负载I/O干扰仍然可以显着扩大I/O延迟。在云计算环境中最好地利用超低延迟SSD仍然是一个开放的问题。在本文中,我们分析了整个I/O堆栈,并揭示I/O干扰主要归因于SSD设备中的资源争夺,交易在文件系统中提交以及昂贵的过程计划。为了解决这些问题,我们提出了Fastresponse,这是一种使用超低延迟SSD进行延迟敏感工作负载的整体方法。首先,我们在块层上提出了一个新的I/O调度程序,以进行油门I/O的请求,以吞吐量为导向的工作负载,从而减少SSD设备中的资源争议。其次,我们开发了一种细粒度的日记帐方案,以减少文件系统层的交易延迟。第三,我们重新设计了完全公平的调度程序(CFS),以促进对延迟敏感过程的优先级。我们在Linux内核中实现Fastresponse,并通过几个混合工作负载对其进行评估。与Vanilla Linux和最先进的SelectiSR相比,Fastresponse可以将对潜伏期敏感工作负载的平均响应时间分别减少18---70%和10---67%,并将99.9%的响应时间分别减少58-------80%和52%和52-----78%。同时,面向吞吐量的工作负载的性能降解小于6%。
Low-latency I/O services are essential for latency-sensitive workloads when they co-run with throughput-oriented workloads in cloud data centers. Although advanced SSDs such as Intel Optane SSDs can offer ultra-low latency at the device layer, I/O interference among various workloads through the I/O stack can still significantly enlarge I/O latency. It is still an open problem to best utilize ultra-low latency SSDs in cloud computing environments. In this paper, we analyze the entire I/O stack and reveal that I/O interference is mainly attributed to resource contention in the SSD device, transactions commit in the file system, and costly process scheduling. To address these problems, we propose FastResponse, a holistic approach to use ultra-low latency SSDs for latency-sensitive workloads. First, we propose a new I/O scheduler at the block layer to throttle I/O requests of throughput-oriented workloads, and thus reduce the resource contention in the SSD device. Second, we develop a fine-grained journaling scheme to reduce the latency of transaction at the file system layer. Third, we redesign Completely Fair Scheduler (CFS) to promote the priority of latency-sensitive processes. We implement FastResponse in Linux kernel and evaluate it with several mixed workloads. Compared with the vanilla Linux and the state-of-the-art SelectISR, FastResponse can reduce the average response time of latency-sensitive workloads by 18--70% and 10--67%, respectively, and reduce the 99.9th percentile response time by 58--80% and 52--78%, respectively. Meanwhile, the performance degradation for throughput-oriented workloads is less than 6%.