Affinity-Based Network Interfaces for Efficient Communication on Multicore Architectures

Affinity-Based Network Interfaces for Efficient Communication on Multicore Architectures
复制标题

基于亲和力的网络接口可在多核架构上实现高效通信

DOI:
10.1007/s11390-013-1352-2
复制
发表时间:
2013
影响因子:
1.9
通讯作者:
A. Prieto
A. Prieto
中科院分区:
计算机科学3区
文献类型:
--
作者:
Andrés Ortiz;Julio Ortega Lopera;A. F. Díaz;A. Prieto

文献摘要

被引文献

相似文献

具有高通信需求的应用程序(例如,一些多媒体、实时和高性能计算应用程序)的需求需要改进网络接口性能,并且网络链路的可用性提供了每秒数千兆比特的带宽,这可能需要许多处理器周期来执行通信任务。多核架构是当前微处理器发展的趋势,以应对进一步提高时钟频率和微架构效率的困难,为利用节点中的并行性来设计高效的通信架构提供了新的机会。然而,尽管目前的操作系统网络堆栈包括多个线程,使得在内核中并发执行网络任务成为可能,但是基于数据包或基于连接的并行性的实现并不是微不足道的,因为它们必须考虑到访问共享资源时同步的成本和缓存的有效使用。因此,将网络中断以及相应的协议和网络应用程序处理分配到相同的核心是近年来该主题研究的一个共同趋势,因为这种亲和调度可以减少对共享资源的争用和缓存丢失。在本文中,我们提出并分析了几种配置,以在服务器中可用的不同内核之间分配网络接口。这些备选方案是根据相应的通信任务与位置(接近存储不同数据结构的存储器)的亲和性和处理核心的特性设计的。由于这种方法使用多个核心来加速给定连接的通信路径,因此可以将其视为对那些考虑多个核心同时处理属于相同或不同连接的数据包的方法的补充。消息传递接口(MPI)工作负载和动态web服务器已被视为评估和比较这些备选方案的通信性能的应用程序。在我们的实验中,通过全系统模拟执行,在MPI工作负载中观察到吞吐量提高了35%,延迟提高了23%,在动态web服务器中测量到吞吐量提高了100%,响应时间提高了500%,每秒处理的请求提高了82%。
Improving the network interface performance is needed by the demand of applications with high communication requirements (for example, some multimedia, real-time, and high-performance computing applications), and the availability of network links providing multiple gigabits per second bandwidths that could require many processor cycles for communication tasks. Multicore architectures, the current trend in the microprocessor development to cope with the difficulties to further increase clock frequencies and microarchitecture efficiencies, provide new opportunities to exploit the parallelism available in the nodes for designing efficient communication architectures. Nevertheless, although present OS network stacks include multiple threads that make it possible to execute network tasks concurrently in the kernel, the implementations of packet-based or connection-based parallelism are not trivial as they have to take into account issues related with the cost of synchronization in the access to shared resources and the efficient use of caches. Therefore, a common trend in many recent researches on this topic is to assign network interrupts and the corresponding protocol and network application processing to the same core, as with this affinity scheduling it would be possible to reduce the contention for shared resources and the cache misses. In this paper we propose and analyze several configurations to distribute the network interface among the different cores available in the server. These alternatives have been devised according to the affinity of the corresponding communication tasks with the location (proximity to the memories where the different data structures are stored) and characteristics of the processing core. As this approach uses several cores to accelerate the communication path of a given connection, it can be seen as complementary to those that consider several cores to simultaneously process packets belonging to either the same or different connections. Message passing interface (MPI) workloads and dynamic web servers have been considered as applications to evaluate and compare the communication performance of these alternatives. In our experiments, performed by full-system simulation, improvements of up to 35% in the throughput and up to 23% in the latency have been observed in MPI workloads, and up to 100% in the throughput, up to 500% in the response time, and up to 82% in the requests attended per second have been measured in dynamic web servers.