HyperData: A Data Transfer Accelerator for Software Data Planes Based on Targeted Prefetching

HyperData: A Data Transfer Accelerator for Software Data Planes Based on Targeted Prefetching
复制标题

HyperData:基于目标预取的软件数据平面数据传输加速器

DOI:
10.1109/iccd53106.2021.00059
复制
发表时间:
2021
期刊:
2021 IEEE 39th International Conference on Computer Design (ICCD)
影响因子:
--
通讯作者:
T. Wenisch
T. Wenisch
中科院分区:
--
文献类型:
--
作者:
Hossein Golestani;T. Wenisch

文献摘要

被引文献

相似文献

数据中心系统依赖于快速,高效的I/O软件堆栈(software数据平面(SDP))在众多流程(或VMS)和I/O设备(NICS,SSD等)之间经常互动。当今的I/O设备和μS规模计算的速度不断增长,因此SDP在整体系统性能和效率中起着至关重要的作用。 SDP,I/O设备和应用程序/VM之间的数据传输是通过共享存储器排列的SDP系统中的HyperData加速器(例如网络数据包或存储块)中的。或者,如今(共享)级别的高速缓存,感谢Intel DDIO等技术。缺乏适当的到达通知机制,而Hyperdata的复杂访问模式则是为了执行靶向预取,其中确切的数据项(或所需的子集)被预取到消费者核心的L1缓存适用于Core-Device和Core-core数据通信,并支持复杂的队列格式,例如Virtio和Multi-Concumer队列。预摘要,发布预取请求和系统级监视集,该集合会导致排队数据到达,并触发预摘要操作。 ART SDP,每核高度只有几百个字节。
Datacenter systems rely on fast, efficient I/O soft-ware stacks—Software Data Planes (SDPs)—to coordinate frequent interaction among myriad processes (or VMs) and I/O devices (NICs, SSDs, etc.). Thanks to the impressive and ever-growing speed of today’s I/O devices and μs-scale computation due to hyper-tenancy and microservice-based applications, SDPs play a crucial role in overall system performance and efficiency. In this work, we aim to enhance data transfer among the SDP, I/O devices, and applications/VMs by designing the HyperData accelerator. Data items in SDP systems, such as network packets or storage blocks, are transferred through shared memory queues. Consumer cores typically access the data from DRAM or, thanks to technologies like Intel DDIO, from the (shared) last-level cache. Today, consumers cannot effectively prefetch such data to nearer caches due to the lack of a proper arrival notification mechanism and the complex access pattern of data buffers. HyperData is designed to perform targeted prefetching, wherein the exact data items (or a required subset) are prefetched to the L1 cache of the consumer core. Furthermore, HyperData is applicable to both core–device and core–core data communication, and it supports complex queue formats like Virtio and multi-consumer queues. HyperData is realized with a per-core programmable prefetcher, which issues the prefetch requests, and a system-level monitoring set, which monitors queues for data arrival and triggers prefetch operations. We show that HyperData improves processing latency by 1.20-2.42× in a simulation of a state-of-the-art SDP, with only a few hundred bytes of per-core overhead.