Floem: A Programming System for NIC-Accelerated Network Applications

Floem: A Programming System for NIC-Accelerated Network Applications
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
--
影响因子:
--
通讯作者:
P. Phothilimthana;Ming Liu;Antoine Kaufmann;Simon Peter;Rastislav Bodík;T. Anderson
P. Phothilimthana;Ming Liu;Antoine Kaufmann;Simon Peter;Rastislav Bodík;T. Anderson
中科院分区:
其他
文献类型:
--
作者:
P. Phothilimthana;Ming Liu;Antoine Kaufmann;Simon Peter;Rastislav Bodík;T. Anderson

文献摘要

被引文献

相似文献

开发将计算卸载到NIC加速器的服务器应用程序是复杂而费力的。开发人员必须探索设计空间,其中包括不同卸载策略的语义变化,以及并行化,程序到资源映射和跨设备程序组件的通信策略的变化。因此,我们设计FLOEM -一种语言,编译器和运行时-用于编程NIC加速应用程序。FLOEM通过提供编程抽象来实现卸载设计探索,以将计算分配给硬件资源;控制逻辑队列到物理队列的映射;访问数据包的字段及其元数据,而无需手动编组数据包;使用NIC来存储昂贵的计算;以及与外部应用程序接口。编译器推断哪些数据必须在CPU和NIC之间传输,并生成完整的缓存实现,而运行时则透明地优化DMA吞吐量。我们使用FLOEM来探索真实世界应用程序的NIC卸载设计,包括键值存储和分布式实时数据分析系统;与仅CPU实现相比,它们的吞吐量分别提高了1.3-3.6倍和75- 96%。
Developing server applications that offload computation to a NIC accelerator is complex and laborious. Developers have to explore the design space, which includes semantic changes for different offloading strategies, as well as variations on parallelization, program-to-resource mapping, and communication strategies for program components across devices. We therefore design FLOEM -- a language, compiler, and runtime -- for programming NIC-accelerated applications. FLOEM enables offload design exploration by providing programming abstractions to assign computation to hardware resources; control mapping of logical queues to physical queues; access fields of a packet and its metadata without manually marshaling a packet; use a NIC to memoize expensive computation; and interface with an external application. The compiler infers which data must be transferred between the CPU and NIC and generates a complete cache implementation, while the runtime transparently optimizes DMA throughput. We use FLOEM to explore NIC-offloading designs of real-world applications, including a key-value store and a distributed real-time data analytics system; improve their throughput by 1.3-3.6× and by 75-96%, respectively, over a CPU-only implementation.