Application-specific Network-on-Chip architecture synthesis based on set partitions and Steiner Trees

Application-specific Network-on-Chip architecture synthesis based on set partitions and Steiner Trees
复制标题

DOI:
10.1109/aspdac.2008.4483955
复制
发表时间:
2008-01
期刊:
2008 Asia and South Pacific Design Automation Conference
影响因子:
--
通讯作者:
Shan Yan;Bill Lin
Shan Yan;Bill Lin
中科院分区:
其他
文献类型:
--
作者:
Shan Yan;Bill Lin

文献摘要

被引文献

相似文献

本文考虑了合成特定应用的片上网络(NoC)架构的问题。我们提出了两个启发式算法,称为聚类和分解,可以系统地检查不同的集合分区的通信流,我们提出了基于矩形斯坦纳树(Rectilinear-Steiner-tree)的算法生成一个有效的网络拓扑结构中的每个组的分区。不同的评估功能,以配合实现后端和相应的实现技术,可以纳入我们的解决方案框架,以评估的实施成本的集合分区和拓扑结构生成。特别是,我们实验了基于70 nm工艺技术的功耗参数的实施成本模型,其中泄漏功率是能源消耗的主要来源。各种NoC基准测试的实验结果表明,我们的合成结果可以平均实现6.92倍的功耗降低最好的标准网格实现。为了进一步衡量我们的启发式算法的有效性,我们还实现了一个精确的算法,枚举所有不同的集合分区。对于可以获得精确结果的基准测试,我们的CLUSTER和DECOMPOSE算法平均可以实现精确结果的1%和2%以内的结果,执行时间都在1秒以内,而精确算法需要多达4.5小时。
This paper considers the problem of synthesizing application-specific network-on-chip (NoC) architectures. We propose two heuristic algorithms called CLUSTER and DECOMPOSE that can systematically examine different set partitions of communication flows, and we propose Rectilinear-Steiner-tree (RST) based algorithms for generating an efficient network topology for each group in the partition. Different evaluation functions in fitting with the implementation backend and the corresponding implementation technology can be incorporated into our solution framework to evaluate the implementation cost of the set partitions and RST topologies generated. In particular, we experimented with an implementation cost model based on the power consumption parameters of a 70 nm process technology where leakage power is a major source of energy consumption. Experimental results on a variety of NoC benchmarks showed that our synthesis results can on average achieve a 6.92 x reduction in power consumption over the best standard mesh implementation. To further gauge the effectiveness of our heuristic algorithms, we also implemented an exact algorithm that enumerates all distinct set partitions. For the benchmarks where exact results could be obtained, our CLUSTER and DECOMPOSE algorithms on average can achieve results within 1% and 2% of exact results, with execution times all under 1 second whereas the exact algorithms took as much as 4.5 hours.