Overcoming the limitations of conventional vector processors

Overcoming the limitations of conventional vector processors
复制标题

克服传统矢量处理器的局限性

DOI:
--
复制
发表时间:
2003
期刊:
30th Annual International Symposium on Computer Architecture, 2003. Proceedings.
影响因子:
--
通讯作者:
D. Patterson
D. Patterson
中科院分区:
--
文献类型:
--
作者:
Christos Kozyrakis;D. Patterson

文献摘要

被引文献

相似文献

尽管矢量处理器在多媒体应用中具有优越的性能,但它们有三个限制,阻碍了它们的广泛接受。首先,集中式矢量寄存器文件的复杂性和大小限制了功能单元的数量。其次,向量指令的精确异常很难实现。第三,矢量处理器需要昂贵的片上存储系统,以支持低访问延迟的高带宽。我们介绍CODE,一个可扩展的矢量微架构,解决了这三个缺点。它是围绕一个聚集的矢量寄存器文件设计的,并使用一个单独的网络进行跨功能单元的操作数传输。通过广泛使用解耦,它可以隐藏跨功能单元的通信延迟,并比集中式组织提供26%的性能改进。CODE可以有效地扩展到8个功能单元,而不需要广泛的指令发布功能。重命名表使群集寄存器文件在指令集级别透明。重命名还可以在性能损失小于5%的情况下为矢量指令提供精确的异常。最后,解耦允许CODE在不使用片上缓存的情况下容忍在亚线性性能下降时内存延迟的大幅增加。因此,CODE可以使用经济的、片外的存储系统。
Despite their superior performance for multimedia applications, vector processors have three limitations that hinder their widespread acceptance. First, the complexity and size of the centralized vector register file limits the number of functional units. Second, precise exceptions for vector instructions are difficult to implement. Third, vector processors require an expensive on-chip memory system that supports high bandwidth at low access latency. We introduce CODE, a scalable vector microarchitecture that addresses these three shortcomings. It is designed around a clustered vector register file and uses a separate network for operand transfers across functional units. With extensive use of decoupling, it can hide the latency of communication across functional units and provides 26% performance improvement over a centralized organization. CODE scales efficiently to 8 functional units without requiring wide instruction issue capabilities. A renaming table makes the clustered register file transparent at the instruction set level. Renaming also enables precise exceptions for vector instructions at a performance loss of less than 5%. Finally, decoupling allows CODE to tolerate large increases in memory latency at sublinear performance degradation without using on-chip caches. Thus, CODE can use economical, off-chip, memory systems.