A SYSTOLIC BLOCK-JACOBI SVD SOLVER FOR PROCESSOR MESHES

A SYSTOLIC BLOCK-JACOBI SVD SOLVER FOR PROCESSOR MESHES
复制标题

处理器网格的脉动块雅可比 SVD 求解器

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
VAJTERSˇIC
VAJTERSˇIC
中科院分区:
--
文献类型:
--
作者:
Gabriel OKSˇA;MARIA´N;VAJTERSˇIC

文献摘要

被引文献

相似文献

对矩阵A [ R m ∈ n,m $ n,m,n偶数]的奇异值分解(SVD)设计了双边块Jacobi算法的静态版本.该算法涉及p处理机上二维环形网格上的并行排序类CO,数学反演算法基于局部数据矩阵的QR分解(QRD)和对角网格处理机上局部SVD的三角Kogbetliantz算法(TKA)。需要对角和非对角网格处理器中局部矩阵的后续更新。我们表明,所有的更新可以实现正交修改的艾德Givens旋转。这些旋转可以通过环形网格有效地并行流水线化在并行处理器的水平环和垂直环中。我们的解决方案要求,每一个网格处理器,O/2 μ m × n × 2 = p(cid:6)脉动处理元件(PE)和额外的延迟元件。时间复杂度可以被估计为Tp <0/2 m/n 3 = 2 = p 1 = 4 n/w D(cid:6),其中w是双侧块雅可比算法中的全局扫描的数量,并且D是全局同步时间步长的长度。每个网格处理器的VLSI面积,由其构造所需的垂直和水平线的长度来测量,可以估计为A < O ½ m n 2 = p(cid:6);每个网格处理器的组合VLSI面积-时间复杂度为AT 2 < O ½ m n 2 m 2 n 3 = p 5 = 4 w 2 D 2(cid:6):理论上的加速比可以估计为Sp < O p 1 = 4 m 2 n = 0 p 1 = 4 m n 3 = 2 n:使用固定内部大小的网格处理器^ m ^ n ;甚至是,可以构造二维超环面网格,并计算矩阵A的奇异值分解,其大小与网格处理器的形状相匹配,即,m = n/4 ^ m = n:在这个意义上,心脏收缩算法是可缩放的。
Wedesignthesystolicversionofthetwo-sidedblock-Jacobialgorithmforthesingularvalue decomposition(SVD)of matrix A [ R m £ n , m $ n and m , n even. The algorithm involves the class CO of parallel orderings on the two-dimensionaltoroidalmeshwith p processors.ThemathematicalbackgroundisbasedontheQRdecomposition(QRD) of local data matrices and on the triangular Kogbetliantz algorithm (TKA) for local SVDs in the diagonal mesh processors. Subsequentupdatesoflocalmatrices inthediagonalas wellasnondiagonalmeshprocessorsare required. We show that all updates can be realized by orthogonal modified Givens rotations. These rotations can be efficiently pipelined in parallel in the horizontal and vertical rings of ffiffiffi p p processor through the toroidal mesh. Our solution requires, per one mesh processor, O ½ð m þ n Þ 2 = p (cid:6) systolic processing elements (PEs) and additional delay elements. The time complexity can be estimated as T p < O ½ð m þ ð n 3 = 2 = p 1 = 4 ÞÞ w D (cid:6) where w is the number of global sweeps in the two-sided block-Jacobi algorithm and D is the length of the global synchronization time step. The VLSI area per meshprocessor,measuredbythenumberofverticalandhorizontalwiresrequiredforitsconstruction,canbeestimatedas A < O ½ð m þ n Þ 2 = p (cid:6) ; and the combined VLSI area–time complexity per mesh processor is AT 2 < O ½ðð m þ n Þ 2 m 2 n 3 = p 5 = 4 Þ w 2 D 2 (cid:6) : The theoretical speedup can be estimated as S p < O ð p 1 = 4 m 2 n = ð p 1 = 4 m þ n 3 = 2 ÞÞ : Using the mesh processorsoffixedinnersize ^ m £ ^ n ; ^ m , ^ n even,itispossibletoconstructthesquaretwo-dimensionaltoroidalmeshand to compute the SVD of matrix A , the size of the which matches the shape of mesh processors, i.e. m = n ¼ ^ m = ^ n : In this sense, the systolic algorithm is scalable.