The effects of memory latency and fine-grain parallelism on Astronautics ZS-1 performance

The effects of memory latency and fine-grain parallelism on Astronautics ZS-1 performance
复制标题

内存延迟和细粒度并行性对 Astronautics ZS-1 性能的影响

DOI:
--
复制
发表时间:
1990
期刊:
Twenty-Third Annual Hawaii International Conference on System Sciences
影响因子:
--
通讯作者:
E. Davidson
E. Davidson
中科院分区:
--
文献类型:
--
作者:
W. Mangione;S. Abraham;E. Davidson

文献摘要

被引文献

相似文献

检查了宇航员ZS-1的性能,即脱钩的访问/执行(DAE)处理器的性能。 CPU由两个子系统组成:一个访问处理器,可处理地址生成和定点操作;和执行处理器,该处理器处理浮点操作。这两个系统通过队列网络进行通信,并以相当脱钩的方式进行操作。该体系结构表现出一种称为Slip的细粒平行性形式,可改善性能。开发了ZS-1的一些性能范围。简单的资源使用计数足以为大多数向量循环建立良好的上限。依赖图用于形成另一个对非矢量环特别有用的结合。这种两型模型的内容是编译器特征,例如循环展开和硬件特征,例如内存延迟。该模型应用于前12个Livermore循环,并将其与各种存储系统的仿真结果进行了比较。此比较表明ZS-1的可耐受性延迟程度与滑移的函数相容易增加,并提供了有关应用程序代码,架构和编译器功能的见解。<< ETX >>
The performance of the Astronautics ZS-1, a decoupled access/execute (DAE) processor, is examined. The CPU is composed of two subsystems: an access processor, which handles address generation and fixed-point operations; and an execute processor, which handles floating-point operations. These two systems communicate through a network of queues and operate in a fairly decoupled manner. This architecture exhibits a form of fine-grain parallelism, called slip, that improves performance. Some performance bounds for the ZS-1 are developed. A simple count of resource usage is sufficient to establish a good upper bound on performance for most vector loops. A dependence graph is used to form another bound that is particularly useful for nonvector loops. This two-bound model accounts for compiler characteristics, such as loop unrolling, and for hardware characteristics, such as memory latency. This model is applied to the first 12 Livermore loops and compared to simulation results for a variety of memory systems. This comparison indicates how well the ZS-1 tolerates increased memory latency as a function of slip and provides insights regarding application codes, architectures, and compiler capabilities.<<ETX>>