Fault-tolerant matrix arithmetic and signal processing on highly concurrent computing structures

Fault-tolerant matrix arithmetic and signal processing on highly concurrent computing structures
复制标题

高并发计算结构上的容错矩阵运算和信号处理

DOI:
10.1109/proc.1986.13535
复制
发表时间:
1986
影响因子:
20.6
通讯作者:
J. Abraham
J. Abraham
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jing;J. Abraham

文献摘要

被引文献

相似文献

在许多实时和科学应用中,高速执行矩阵运算和信号处理算法的硬件需求量很大。随着超大规模集成电路技术的出现,以高速相互协作的大量处理元件在经济上已经变得可行。由于高性能系统中的任何功能错误都可能严重危及系统的运行及其数据完整性,因此必须纳入一定程度的容错,以确保长时间计算的结果有效。由于许多重要的实时信号处理任务的主要计算要求可以减少到一个共同的基本矩阵运算,矩阵运算的统一容错计划的发展可以解决可靠的信号处理和可靠的矩阵运算的问题。早期的工作提出了一个低成本的校验和计划的容错矩阵操作多处理器系统。然而,该方案只能纠正矩阵乘法中的错误;它可以检测但不能纠正矩阵向量乘法、LU分解、矩阵求逆等中的错误。为了解决校验和方案的这些问题,本文提出了一种非常通用的矩阵编码方案,以实现线性阵列的容错矩阵运算和信号处理,其被认为在VLSI计算结构中具有最大的希望,因为它们的灵活性、低成本和对大多数感兴趣的算法的适用性。因此,提出的技术是一个非常具有成本效益的编码技术,以实现容错矩阵运算和信号处理的高度并发的VLSI计算结构。
Hardware for executing matrix arithmetic and signal processing algorithms at high speeds is in great demand in many real-time and scientific applications. With the advent of VLSI technology, large numbers of processing elements which cooperate with each other at high speed have become economically feasible. Since any functional error in a high-performance system may seriously jeopardize the operation of the system and its data integrity, some level of fault tolerance must be incorporated in order to ensure that the results of long computations are valid. Since the major computational requirements for many important real-time signal processing tasks can be reduced to a common set of basic matrix operations, the development of a unified fault-tolerant scheme for matrix operations can solve the problems of both reliable signal processing and reliable matrix operations. Earlier work proposed a low-cost checksum scheme for fault-tolerant matrix operations on multiple processor systems. However, this scheme can only correct errors in matrix multiplication; it can detect, but not correct, errors in matrix-vector multiplication, LU decomposition, matrix inversion, etc. In order to solve these problems with the checksum scheme, a very general matrix encoding scheme is proposed in this paper to achieve fault-tolerant matrix arithmetic and signal processing with linear arrays, which are believed to hold the most promise in VLSI computing structures for their flexibility, low cost, and applicability to most of the interesting algorithms. This proposed technique is, therefore, a very cost-effective encoding technique to achieve fault-tolerant matrix arithmetic and signal processing on highly concurrent VLSI computing structures.