Tarantula: a vector extension to the alpha architecture

Tarantula: a vector extension to the alpha architecture
复制标题

Tarantula:对 alpha 架构的向量扩展

DOI:
10.1109/isca.2002.1003586
复制
发表时间:
2002
期刊:
Proceedings 29th Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
André Seznec
André Seznec
中科院分区:
--
文献类型:
--
作者:
R. Espasa;Federico Ardanaz;J. Gago;R. Gramunt;I. Hernandez;Toni Juan;J. Emer;S. Felix;P. G. Lowney;Matthew Mattina;André Seznec

文献摘要

被引文献

相似文献

Tarantula是针对技术,科学和生物信息学工作负载的侵略性浮点机,最初计划作为EV8处理器的后续候选人。 Tarantula为EV8核心增加了一个矢量单位,每个周期能够为32个双精度拖鞋。向量单元直接从16 MBYTE的第二级缓存中获取数据,其峰带宽为每个周期64个64位值。整个芯片都由能够输送超过64 gbytes/s的原始带宽的内存控制器支持。 Tarantula通过在新的建筑状态下运行的新向量说明扩展了Alpha ISA。体系结构和实现的显着特征是:(1)它完全集成到虚拟内存缓存系统中,而无需更改其相干协议,(2)为非单位步幅内存访问提供了很高的带宽,(3)支持收集的支持收集。 /散射指令有效地与EV8核心完全集成,具有狭窄的简化界面,而不是充当处理器(5)可以实现每个周期104个操作的峰值,并且(6)实现了真实的“真实” - 计算“每个晶体管和每瓦比率。我们的详细模拟表明,在8倍方面,狼蛛的平均速度超过了EV8的平均速度为5倍。此外,在聚集/分散的基准测试(例如radix排序)上的性能也非常出色:每周周期的EV8和15次持续操作的速度几乎为3倍。每个周期的几个基准超过20个操作。
Tarantula is an aggressive floating point machine targeted at technical, scientific and bioinformatics workloads, originally planned as a follow-on candidate to the EV8 processor. Tarantula adds to the EV8 core a vector unit capable of 32 double-precision flops per cycle. The vector unit fetches data directly from a 16 MByte second level cache with a peak bandwidth of sixty four 64-bit values per cycle. The whole chip is backed by a memory controller capable of delivering over 64 GBytes/s of raw bandwidth. Tarantula extends the Alpha ISA with new vector instructions that operate on new architectural state. Salient features of the architecture and implementation are: (1) it fully integrates into a virtual-memory cache-coherent system without changes to its coherency protocol, (2) provides high bandwidth for non-unit stride memory accesses, (3) supports gather/scatter instructions efficiently, (4) fully integrates with the EV8 core with a narrow, streamlined interface, rather than acting as a co-processor (5) can achieve a peak of 104 operations per cycle, and (6) achieves excellent "real-computation" per transistor and per watt ratios. Our detailed simulations show that Tarantula achieves an average speedup of 5X over EV8, out of a peak speedup in terms of flops of 8X. Furthermore, performance on gather/scatter intensive benchmarks such as Radix Sort is also remarkable: a speedup of almost 3X over EV8 and 15 sustained operations per cycle. Several benchmarks exceed 20 operations per cycle.