SRC: Mitigate I/O Throughput Degradation in Network Congestion Control of Disaggregated Storage Systems

SRC: Mitigate I/O Throughput Degradation in Network Congestion Control of Disaggregated Storage Systems
复制标题

DOI:
10.1109/ipdps54959.2023.00035
复制
发表时间:
2023-05
期刊:
2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Danlin Jia;Yiming Xie;Li Wang;Xiaoqian Zhang;Allen Yang;Xuebin Yao;Mahsa Bayati;Pradeep Subedi;B. Sheng;N. Mi
Danlin Jia;Yiming Xie;Li Wang;Xiaoqian Zhang;Allen Yang;Xuebin Yao;Mahsa Bayati;Pradeep Subedi;B. Sheng;N. Mi
中科院分区:
其他
文献类型:
--
作者:
Danlin Jia;Yiming Xie;Li Wang;Xiaoqian Zhang;Allen Yang;Xuebin Yao;Mahsa Bayati;Pradeep Subedi;B. Sheng;N. Mi

文献摘要

相似文献

业界已采用分散式存储系统为超大规模架构提供高质量的服务。此基础架构使组织能够访问可独立管理、配置和扩展的存储资源。它得到了全闪存阵列和NVMe-over-Fabric协议的最新进展的支持,从而能够通过不同的网络结构远程访问NVMe设备。针对传统的远程直接内存访问协议(RDMA)中的网络拥塞问题,提出了一种新的解决方案。然而,NVMe-oF提出了新的挑战,在拥塞控制的分散存储systems.In这项工作中,我们调查的性能下降的读吞吐量的存储节点所造成的传统的网络拥塞控制机制。我们设计了一个存储端速率控制(SRC),以缓解网络拥塞,同时避免存储节点的性能下降。首先,我们在NVMe驱动程序层设计了一个I/O吞吐量控制机制,以实现对存储节点的吞吐量控制。其次,我们构建了一个吞吐量预测模型,学习工作负载特性和I/O吞吐量之间的映射函数。第三,我们在存储节点上部署SRC,以配合NVMe-over-RDMA架构上的传统网络拥塞控制。最后,我们评估SRC与不同的工作负载,SSD配置和网络拓扑。实验结果表明,SRC取得了显着的性能改善。
The industry has adopted disaggregated storage systems to provide high-quality services for hyper-scale architectures. This infrastructure enables organizations to access storage resources that can be independently managed, configured, and scaled. It is supported by the recent advances of all-flash arrays and NVMe-over-Fabric protocol, enabling remote access to NVMe devices over different network fabrics. A surge of research has been proposed to mitigate network congestion in traditional remote direct memory access protocol (RDMA). However, NVMe-oF raises new challenges in congestion control for disaggregated storage systems.In this work, we investigate the performance degradation of the read throughput on storage nodes caused by traditional network congestion control mechanisms. We design a storage-side rate control (SRC) to relieve network congestion while avoiding performance degradation on storage nodes. First, we design an I/O throughput control mechanism in the NVMe driver layer to enable throughput control on storage nodes. Second, we construct a throughput prediction model to learn a mapping function between workload characteristics and I/O throughput. Third, we deploy SRC on storage nodes to cooperate with traditional network congestion control on an NVMe-over-RDMA architecture. Finally, we evaluate SRC with varying workloads, SSD configurations, and network topologies. The experimental results show that SRC achieves significant performance improvement.