Performance modeling of multiprocessor implementations of protocols

Performance modeling of multiprocessor implementations of protocols
复制标题

协议的多处理器实现的性能建模

DOI:
10.1109/90.700890
复制
发表时间:
1998
期刊:
IEEE/ACM Trans. Netw.
影响因子:
--
通讯作者:
P. Gunningberg
P. Gunningberg
中科院分区:
--
文献类型:
--
作者:
M. Björkman;P. Gunningberg

文献摘要

被引文献

相似文献

协议的多处理器执行中的两个主要性能瓶颈是共享内存和锁的争用。锁用于保护由竞争处理器共享的存储器中的共享消息和/或共享协议状态。如果并行协议代码频繁地访问共享状态和数据,则在锁争用和存储器争用方面,通过锁定的互斥可能是昂贵的。本文提出了一种用于预测共享内存多处理器协议执行性能的嵌入式网络模型。从这个模型的预测进行比较,从两个常用的通信协议栈,传输控制协议/互联网协议(TCP/IP)/以太网和用户数据报协议/互联网协议(UDP/IP)/以太网的多处理器实现的性能测量。这些栈是在亚利桑那大学的x-kernel协议环境的并行化版本上实现的。一个“处理器每消息”的范例是用来划分负载之间的处理器。并行实现相对于顺序实现的加速比对于UDP(使用20个处理器)是11倍以上,对于TCP(使用5个处理器)是3倍以上。我们表明,该模型准确地捕捉锁和内存争用在我们的共享内存多处理器的影响,并预测的性能差异小于10%。
Two major performance bottlenecks in multiprocessor execution of protocols are contention for shared memory and for locks. Locks are used to protect shared messages and/or shared protocol state in a memory shared by competing processors. Mutual exclusion by locking can be costly, in terms of both lock contention and memory contention, if the parallel protocol code frequently accesses shared state and data. This paper presents a queueing network model for performance predictions of shared-memory multiprocessor protocol executions. Predictions from this model are compared to performance measurements from a multiprocessor implementation of two commonly used communication protocol stacks, transmission control protocol/Internet protocol (TCP/IP)/Ethernet and user datagram protocol/Internet protocol (UDP/IP)/Ethernet. These stacks are implemented on a parallelized version of the x-kernel protocol environment from the University of Arizona. A "processor-per-message" paradigm is used to partition the load among the processors. The measured speedups for the parallel implementations relative to the sequential ones are more than 11 times for UDP (using 20 processors) and three times for TCP (using five processors) on a sequent symmetry. We show that the model accurately captures the effects of lock and memory contention in our shared-memory multiprocessor and predicts the performance with a discrepancy of less than 10%.