Understanding host network stack overheads

Understanding host network stack overheads
复制标题

DOI:
10.1145/3452296.3472888
复制
发表时间:
2021-08
期刊:
Proceedings of the 2021 ACM SIGCOMM 2021 Conference
影响因子:
--
通讯作者:
Qizhe Cai;Shubham Chaudhary;Midhul Vuppalapati;Jaehyun Hwang;R. Agarwal
Qizhe Cai;Shubham Chaudhary;Midhul Vuppalapati;Jaehyun Hwang;R. Agarwal
中科院分区:
其他
文献类型:
--
作者:
Qizhe Cai;Shubham Chaudhary;Midhul Vuppalapati;Jaehyun Hwang;R. Agarwal

文献摘要

相似文献

传统的终端主机网络栈由于其不可持续的CPU开销,难以跟上数据中心接入链路带宽的快速增长。受此驱动,我们的社区正在探索多种面向未来网络栈的解决方案:从Linux内核优化到部分硬件卸载,从全新的用户空间栈到专用的主机网络硬件。这些解决方案所探索的设计空间将受益于对现有网络栈中CPU低效问题的详细了解。本文介绍了针对100Gbps接入链路带宽的Linux内核网络栈性能的测量结果和见解。我们的研究表明,如此高的带宽链路,加上其他主机资源(例如CPU速度和容量、缓存大小、网卡缓冲区大小等)相对停滞的技术趋势,标志着主机网络栈瓶颈的根本性转变。例如,我们发现单个核心不再能够以线速处理数据包,接收端从内核到应用缓冲区的数据复制成为核心性能瓶颈。此外,带宽延迟积的增长超过了缓存大小的增长,导致网卡和CPU之间的直接内存访问(DMA)流水线效率低下。最后,我们发现现有操作系统中网络栈和CPU调度器的传统松散耦合设计成为跨核心扩展网络栈性能的一个限制因素。基于我们研究的见解,我们讨论了对未来操作系统、网络协议和主机硬件设计的影响。
Traditional end-host network stacks are struggling to keep up with rapidly increasing datacenter access link bandwidths due to their unsustainable CPU overheads. Motivated by this, our community is exploring a multitude of solutions for future network stacks: from Linux kernel optimizations to partial hardware offload to clean-slate userspace stacks to specialized host network hardware. The design space explored by these solutions would benefit from a detailed understanding of CPU inefficiencies in existing network stacks. This paper presents measurement and insights for Linux kernel network stack performance for 100Gbps access link bandwidths. Our study reveals that such high bandwidth links, coupled with relatively stagnant technology trends for other host resources (e.g., CPU speeds and capacity, cache sizes, NIC buffer sizes, etc.), mark a fundamental shift in host network stack bottlenecks. For instance, we find that a single core is no longer able to process packets at line rate, with data copy from kernel to application buffers at the receiver becoming the core performance bottleneck. In addition, increase in bandwidth-delay products have outpaced the increase in cache sizes, resulting in inefficient DMA pipeline between the NIC and the CPU. Finally, we find that traditional loosely-coupled design of network stack and CPU schedulers in existing operating systems becomes a limiting factor in scaling network stack performance across cores. Based on insights from our study, we discuss implications to design of future operating systems, network protocols, and host hardware.