Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors

Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors
复制标题

芯片多处理器的编译器辅助数据分发和网络配置

DOI:
--
复制
发表时间:
2012
影响因子:
5.3
通讯作者:
A. Jones
A. Jones
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yong Li;Ahmed Abousamra;R. Melhem;A. Jones

文献摘要

被引文献

相似文献

数据访问延迟是芯片多处理器性能的一个限制因素,随着具有分布式缓存库的非均匀缓存架构中的核心数量的增加而显着增加。为了减轻这种影响,我们使用基于编译器的方法来利用数据访问的本地性,选择一个优化的数据放置和有效地配置片上网络。建议的实验编译器框架采用新的编译技术来发现和表示多线程内存访问模式(MMAPs)。在运行时,符号MMAP被解析并由分区算法使用,以在所分析的应用程序中的分叉线程之间选择分配的内存块的分区。该分区用于通过将数据与执行拥有数据的线程的核心相关联来强制数据所有权。基于该划分,可以提取应用的通信模式。我们演示了如何将这些信息用于实验架构,以加速应用程序。特别是,我们的编译器辅助数据分区方法显示了20%的加速比共享缓存和5%的加速比最接近的运行时近似,第一次触摸。通过利用通信模式,我们可以实现与运行时使用复杂的集中式网络配置系统的系统相当的性能。因此,我们的最终系统节省了显着的运行时的复杂性,并实现了5.1%的额外加速通过添加的可重构网络。
Data access latency, a limiting factor in the performance of chip multiprocessors, grows significantly with the number of cores in nonuniform cache architectures with distributed cache banks. To mitigate this effect, we use a compiler-based approach to leverage data access locality, choose an optimized data placement and efficiently configure the on-chip network. The proposed experimental compiler framework employs novel compilation techniques to discover and represent multithreaded memory access patterns (MMAPs). At runtime, symbolic MMAPs are resolved and used by a partitioning algorithm to choose a partition of allocated memory blocks among the forked threads in the analyzed application. This partition is used to enforce data ownership by associating the data with the core that executes the thread owning the data. Based on the partition, the communication pattern of the application can be extracted. We demonstrate how this information can be used in an experimental architecture to accelerate applications. In particular, our compiler assisted data partitioning approach shows a 20 percent speedup over shared caching and 5 percent speedup over the closest runtime approximation, first touch. By leveraging the communication pattern we can achieve a comparable performance to a system that uses a complex centralized network configuration system at runtime. Thus, our final system saves significant runtime complexity and achieves an 5.1 percent additional speedup through the addition of the reconfigurable network.