On the Parallelization of Vector Fitting Algorithms

On the Parallelization of Vector Fitting Algorithms
复制标题

DOI:
10.1109/tcpmt.2011.2167973
复制
发表时间:
2011-11-01
影响因子:
2.2
通讯作者:
Grivet-Talocia, Stefano
Grivet-Talocia, Stefano
中科院分区:
工程技术3区
文献类型:
--
作者:
Chinea, Alessandro;Grivet-Talocia, Stefano

文献摘要

被引文献

相似文献

所谓的矢量拟合(VF)算法在过去几年中得到了广泛的应用。该技术提供了一种非常有效的系统辨识工具,从线性和时不变系统的输入-输出响应开始,计算其传递矩阵的有理近似。后者通常用于合成紧凑的宽带等效电路或状态空间模型的可能复杂的互连在芯片,封装,板,甚至系统级。VF算法是基于迭代线性最小二乘解和特征解的组合,并证明了鲁棒性和可靠性。VF的一个潜在的弱点是其相对较差的可扩展性与建模下的结构的复杂性。当输入-输出端口的数量非常大时(一百个或更多,如在电源总线或封装的情况下),过多的计算要求可能会阻碍VF性能并阻止其成功应用。在本文中,我们解决这些问题,首先提出了一个详细的分析计算成本的所有算法部分。结果显示了多核硬件VF并行化的巨大潜力,并提出了一些替代并行化策略。每一种策略都有详细的描述。最后,数值结果和比较提供了一个大的工业基准。这些结果表明,该算法的并行部分具有良好的可扩展性和加速因子,从而大大减少了整体运行时间。
The so-called vector fitting (VF) algorithm has gained much popularity over the last few years. This technique provides a very effective system identification tool that, starting from input-output responses of a linear and time-invariant system, computes a rational approximation of its transfer matrix. The latter is routinely used to synthesize compact broadband equivalent circuits or state-space models of possibly complex interconnects at the chip, package, board, or even system level. The VF algorithm is based on a combination of iterative linear least squares solutions and eigensolutions, and proves robust and reliable. A potential weak point of VF is its relatively poor scalability with the complexity of the structure under modeling. When the number of input-output ports is very large (one hundred or more, as in the case of power buses or packages), the excessive computational requirements may hinder VF performance and prevent its successful application. In this paper, we address these issues by first presenting a detailed analysis of the computational cost of all the algorithm parts. The results show a very good potential for VF parallelization for multicore hardware, and suggest a few alternative parallelization strategies. Each of these strategies is described in detail. Finally, numerical results and comparisons are provided on a large set of industrial benchmarks. These results demonstrate excellent scalability and speedup factors for the parallel sections of the algorithm, leading to a drastic reduction in overall runtime.