Scan primitives for vector computers

Scan primitives for vector computers
复制标题

扫描向量计算机的基元

DOI:
10.1109/superc.1990.130084
复制
发表时间:
1990
期刊:
Proceedings SUPERCOMPUTING '90
影响因子:
--
通讯作者:
M. Zagha
M. Zagha
中科院分区:
--
文献类型:
--
作者:
S. Chatterjee;G. Blelloch;M. Zagha

文献摘要

被引文献

相似文献

作者描述了一组在Cray Y-MP的单个处理器上的一组扫描(也称为All-Prefix-sums)的优化实现现有的计算机技术。用于实施扫描的算法基于并行计算机的算法。这些扫描的一组分段版本仅比未分段版本略高。作者根据扫描的扫描速度描述了一个radix排序程序,该扫描速度比Fortran版本快13倍,在高度优化的库排序例程的20%以内,在树上的三个操作比相应的C版本快10到20倍,连接师学习算法比稀疏和不规则网络的相应C版本快10倍。<< etx >>
The authors describe an optimized implementation of a set of scan (also called all-prefix-sums) primitives on a single processor of a CRAY Y-MP, and demonstrate that their use leads to greatly improved performance for several applications that cannot be vectorized with existing computer technology. The algorithm used to implement the scans is based on an algorithm for parallel computers. A set of segmented versions of these scans is only marginally more expensive than the unsegmented versions. The authors describe a radix sorting routine based on the scans that is 13 times faster than a Fortran version and within 20% of a highly optimized library sort routine, three operations on trees that are between 10 to 20 times faster than the corresponding C versions, and a connectionist learning algorithm that is 10 times faster than the corresponding C version for sparse and irregular networks.<<ETX>>