ENERGY EFFICIENT CLUSTER COPROCESSORS

ENERGY EFFICIENT CLUSTER COPROCESSORS
复制标题

高能效集群协处理器

DOI:
--
复制
发表时间:
--
期刊:
--
影响因子:
--
通讯作者:
A. Davis
A. Davis
中科院分区:
--
文献类型:
--
作者:
Ali Ibrahim;Mike Parker;A. Davis

文献摘要

被引文献

相似文献

新的 3G 无线算法需要比当前嵌入式处理器所能提供的更高的性能。 ASIC 提供了必要的性能,但设计成本高昂且牺牲了通用性。本文介绍了一种集群 VLIW 协处理器方法,该方法以不同于传统通用处理器或 DSP 的方式组织执行和存储资源。协处理器的执行单元被集群化并嵌入到丰富的通信资源中。这些资源的细粒度控制是由宽字水平微代码程序强加的。这种方法的优点在一套六种算法上得到了量化,这些算法取自传统的 DSP 应用和新的 3G 蜂窝电话领域。结果令人惊讶。与英特尔 XScale 等传统嵌入式处理器相比,执行集群保留了传统处理器的大部分通用性,同时将性能提高了一到两个数量级,并将能量延迟减少了三到四个数量级。
New 3G wireless algorithms require more performance than can be currently provided by embedded processors. ASICs provide the necessary performance but are costly to design and sacrifice generality. This paper introduces a clustered VLIW coprocessor approach that organizes the execution and storage resources differently than a traditional general-purpose processor or DSP. The execution units of the coprocessor are clustered and embedded in a rich set of communication resources. Fine grain control of these resources is imposed by a wide-word horizontal micro-code program. The advantages of this approach are quantified on a suite of six algorithms that are taken from both traditional DSP applications and from the new 3G cellular telephony domain. The result is surprising. The execution clusters retain much of the generality of a conventional processor while simultaneously improving performance by one to two orders of magnitude and by reducing energy-delay by three to four orders of magnitude when compared to a conventional embedded processor such as the Intel XScale.