Understanding PARSEC performance on contemporary CMPs

Understanding PARSEC performance on contemporary CMPs
复制标题

了解当代 CMP 上的 PARSEC 性能

DOI:
10.1109/iiswc.2009.5306793
复制
发表时间:
2009
期刊:
2009 IEEE International Symposium on Workload Characterization (IISWC)
影响因子:
--
通讯作者:
S. Mckee
S. Mckee
中科院分区:
--
文献类型:
--
作者:
M. Bhadauria;Vincent M. Weaver;S. Mckee

文献摘要

被引文献

相似文献

PARSEC是工业和学术界用于评估新的芯片多处理器(CMP)设计的参考应用程序套件。到目前为止,还没有任何调查对真实硬件上的Parsec进行分析,以更好地了解扩展属性和瓶颈。这一理解对于指导未来针对这些新出现的工作负载的CMP设计至关重要。我们使用硬件性能计数器,采取系统级方法,并改变常见的体系结构参数:无序核心的数量、内存层次结构配置、多个并发线程的数量、内存通道数量和处理器频率。我们发现这些程序在很大程度上是受计算限制的,因此受到内核数量、微体系结构资源和高速缓存到高速缓存传输的限制,而不是芯片外存储器或系统总线带宽。一半的套件无法随着线程数量的增加而线性扩展,一些应用程序在所有测试平台上的少数几个线程上会导致性能饱和。利用线程级并行性比利用指令级并行性带来更大的回报。为了降低功耗和提高性能,我们建议增加每个内核的算术单元数量,增加对TLP的支持,并减少对ILP的支持。
PARSEC is a reference application suite used in industry and academia to assess new Chip Multiprocessor (CMP) designs. No investigation to date has profiled PARSEC on real hardware to better understand scaling properties and bottlenecks. This understanding is crucial in guiding future CMP designs for these kinds of emerging workloads. We use hardware performance counters, taking a systems-level approach and varying common architectural parameters: number of out-of-order cores, memory hierarchy configurations, number of multiple simultaneous threads, number of memory channels, and processor frequencies. We find these programs to be largely compute-bound, and thus limited by number of cores, micro-architectural resources, and cache-to-cache transfers, rather than by off-chip memory or system bus bandwidth. Half the suite fails to scale linearly with increasing number of threads, and some applications saturate performance at few threads on all platforms tested. Exploiting thread level parallelism delivers greater payoffs than exploiting instruction level parallelism. To reduce power and improve performance, we recommend increasing the number of arithmetic units per core, increasing support for TLP, and reducing support for ILP.