A Throughput-Driven Task Creation and Mapping for Network Processors

A Throughput-Driven Task Creation and Mapping for Network Processors
复制标题

网络处理器的吞吐量驱动任务创建和映射

DOI:
--
复制
发表时间:
2007
期刊:
International Conference on High Performance Embedded Architectures and Compilers
影响因子:
--
通讯作者:
R. Ju
R. Ju
中科院分区:
--
文献类型:
--
作者:
Lixia Liu;Xiao;Michael K. Chen;R. Ju

文献摘要

被引文献

相似文献

网络处理器是可以高速处理数据包的可编程设备。网络处理器的特点是多线程和异构多处理,这通常需要程序员手动创建多个任务并将这些任务映射到不同的处理元素上。本文解决了自动创建任务并将网络应用程序映射到底层硬件以最大化其吞吐量的问题。我们提出了一种吞吐量成本模型来指导任务创建和映射,其目标是最小化处理管道中的阶段数量并同时最大化最慢任务的平均吞吐量。平均吞吐量是通过考虑通信成本、计算成本、内存访问延迟和同步成本来建模的。我们设想程序员为网络应用程序编写小函数,以便我们使用分组和复制来从函数构造任务。从 m 个函数创建任务并将其映射到 n 个处理器的最佳解决方案是一个 NP 困难问题。因此,我们提出了一种实用且高效的启发式算法,其复杂度为 O((n+m)m),并表明所获得的解决方案为典型的网络应用提供了出色的性能。整个框架已在开放研究编译器(ORC)中实现,该编译器适合编译以特定领域数据流语言编写的网络应用程序。实验结果表明,我们的编译器生成的代码可以在 OC-48 输入线速率上实现 100% 的吞吐量。 OC-48 是一种光纤连接,可以处理 2.488Gbps 连接速度,这正是我们的目标硬件的设计目的。我们还证明了良好的创建和映射选择对于实现高吞吐量的重要性。此外,我们还表明,降低通信成本和高效的资源管理是最大限度提高 Intel IXP 网络处理器吞吐量的最重要因素。
Network processors are programmable devices that can process packets at a high speed. A network processor is typified by multi-threading and heterogeneous multiprocessing, which usually requires programmers to manually create multiple tasks and map these tasks onto different processing elements. This paper addresses the problem of automating task creation and mapping of network applications onto the underlying hardware to maximize their throughput. We propose a throughput cost model to guide the task creation and mapping with the objective of both minimizing the number of stages in the processing pipeline and maximizing the average throughput of the slowest task simultaneously. The average throughput is modeled by taking communication cost, computation cost, memory access latency and synchronization cost into account. We envision that programmers write small functions for network applications, such that we use grouping and duplication to construct tasks from the functions. The optimal solution of creating tasks from m functions and mapping them to n processors is an NP-hard problem. Therefore, we present a practical and efficient heuristic algorithm with an O((n+m)m) complexity and show that the obtained solutions produce excellent performance for typical network applications. The entire framework has been implemented in the Open Research Compiler (ORC) adapted to compile network applications written in a domain-specific dataflow language. Experimental results show that the code produced by our compiler can achieve the 100% throughput on the OC-48 input line rate. OC-48 is a fiber optic connection that can handle a 2.488Gbps connection speeds, which is what our targeted hardware was designed for. We also demonstrate the importance of good creation and mapping choices on achieving high throughput. Furthermore, we show that reducing communication cost and efficient resource management are the most important factors for maximizing throughput on the Intel IXP network processors.