A High-Performance and Scalable NVMe Controller Featuring Hardware Acceleration

A High-Performance and Scalable NVMe Controller Featuring Hardware Acceleration
复制标题

DOI:
10.1109/tcad.2021.3088784
复制
发表时间:
2022-05
影响因子:
2.9
通讯作者:
Yunhui Qiu;Wenbo Yin;Lingli Wang
Yunhui Qiu;Wenbo Yin;Lingli Wang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yunhui Qiu;Wenbo Yin;Lingli Wang

文献摘要

相似文献

Nonvolatile Memory Express(NVMe)是一种基于PCI Express(PCIe)的高性能和可扩展接口,用于主机软件与NVM通信,包括NAND闪存和存储类存储器(SCM)。NVMe固态硬盘(SSD)已部署在云平台和数据中心中,用于各种I/O密集型应用程序,因为与SATA/SAS SSD相比,它们具有性能优势。考虑到设计灵活性,基于固件的NVMe控制器通常用于基于闪存的NVMe SSD中,但可能占用很大一部分处理器资源和功耗以实现高性能。此外,固件组件可能是比闪存快一个数量级的SCM的关键性能瓶颈。为了应对这些挑战,工业界和学术界都出现了硬件加速的NVMe控制器。商业硬件控制器是保密的,而目前的学术研究仍然为架构创新留出了很大的空间。在本文中,我们提出了一种开源的超低延迟和高吞吐量NVMe控制器,具有高度并行,流水线和可扩展的架构,可容纳一个管理控制器和多个完全硬件自动化的I/O控制器。我们对NVMe I/O大小、队列深度、队列数量、读写比和访问模式进行了广泛的经验性能评估。最大读写带宽可达到7.0 GB/s,占PCIe带宽的89%。4 KB大小的读写吞吐量可以达到每秒170万次I/O操作(MIOPS),而平均延迟仅为2.4 $\mu \text{s}$ /3.2 $\mu \text{s}$。与学术界最先进的NVMe控制器相比,我们的控制器的4 KB大小的读/写带宽高达2.2\times /2.3\times $,延迟低5.1\times /4.9\times $。
Nonvolatile memory express (NVMe) is a high-performance and scalable PCI express (PCIe)-based interface for the host software communicating with NVMs, including NAND Flash and the storage class memories (SCMs). NVMe solid-state drives (SSDs) have been deployed in cloud platforms and data-centers for a variety of I/O intensive applications due to their performance benefits compared to SATA/SAS SSDs. Considering the design flexibility, firmware-based NVMe controllers are typically used in Flash-based NVMe SSDs but may occupy a significant portion of processor resources and power consumption to achieve high performance. Moreover, the firmware component can be a critical performance bottleneck for SCMs that are an order-of-magnitude faster than Flash. To address these challenges, hardware-accelerated NVMe controllers have emerged in both industry and academia. The commercial hardware controllers are confidential, whereas current academic studies still spare much room for architecture innovations. In this article, we propose an opensource ultralow-latency and high-throughput NVMe controller with a highly parallel, pipelined, and scalable architecture that accommodates one admin controller and multiple fully hardware-automated I/O controllers. We perform extensive empirical performance evaluations concerning the NVMe I/O size, queue depth, queue number, read-to-write ratio, and access pattern. The maximum read/write bandwidth can achieve 7.0 GB/s, accounting for 89% of the PCIe bandwidth. The 4-KB-sized read/write throughput can attain 1.7 million I/O operations per second (MIOPS), whereas the average latency is merely 2.4 $\mu \text{s}$ /3.2 $\mu \text{s}$ . Compared to state-of-the-art NVMe controllers in academia, the 4-KB-sized read/write bandwidth of our controller reaches $2.2 \times /2.3\times $ as high and the latency is $5.1 \times /4.9\times $ lower.