Latency Sensitive FMA Design

Latency Sensitive FMA Design
复制标题

延迟敏感的 FMA 设计

DOI:
--
复制
发表时间:
2011
期刊:
IEEE Symposium on Computer Arithmetic
影响因子:
--
通讯作者:
M. Horowitz
M. Horowitz
中科院分区:
--
文献类型:
--
作者:
Sameh Galal;M. Horowitz

文献摘要

被引文献

相似文献

合并浮点乘加运算的实现可以通过多种方式进行优化。对于延迟敏感的应用,我们的级联设计将累积相关延迟降低了2倍,而非累积相关延迟增加了13%。一个简单的按序执行模型表明,这种设计在大多数应用程序中是上级的,提供了12%的FP暂停平均减少,并提高性能高达6%。超标量乱序机器的仿真显示,2路机器的CPI平均提高了4%,4路机器的CPI平均提高了4.6%。级联设计具有与传统的融合多加FMA相同的面积和能量预算。
The implementation of merged floating-point multiply-add operations can be optimized in many ways. For latency sensitive applications, our cascade design reduces the accumulation dependent latency by 2x over a fused design, at a cost of a 13% increase in non-accumulation dependent latency. A simple in-order execution model shows this design is superior in most applications, providing 12% average reduction in FP stalls, and improves performance by up to 6%. Simulations of superscalar out-of-order machines show 4% average improvement in CPI in 2-way machines and 4.6% in 4-way machines. The cascade design has the same area and energy budget as a traditional fused multiple-add FMA.