Matrix product on heterogeneous master-worker platforms

Matrix product on heterogeneous master-worker platforms
复制标题

异构主从平台上的矩阵乘积

DOI:
--
复制
发表时间:
2008
期刊:
ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子:
--
通讯作者:
F. Vivien
F. Vivien
中科院分区:
--
文献类型:
--
作者:
J. Dongarra;J. Pineau;Y. Robert;F. Vivien

文献摘要

被引文献

相似文献

本文的重点是设计有效的平行矩阵 - 产品算法,用于异质的大师级平台。尽管矩阵产物对处理器的均质2D阵列(例如,Cannon算法和Scalapack外部产品算法)都充分理解,但有三个关键的假设使我们的作品具有原创性和创新性: - 集中数据。我们假设所有矩阵文件源于主体,并且必须返回到主。主人在Scalapack,输入和输出矩阵中将数据和计算分配给工人,应该事先将其分配在参与资源之间)。通常,我们的方法在加快在服务器上运行的MATLAB或SCILAB客户端的背景下很有用(充当文件的主和初始存储库)。 - 异质的星形平台。我们针对完全异质的平台,其中计算资源具有不同的计算能力。此外,工人通过不同能力的链接连接到主人。从服务器部署应用程序时,该框架是现实的,该应用程序负责注册授权资源。 - 有限的内存。当我们研究大问题的并行化时,我们不能假设可以将完整的矩阵列块存储在工作人员的记忆中,并重复使用以进行后续更新(如Scalapack中)。我们已经设计了有效的资源选择算法(确定要注册的工人)和通信顺序(用于输入和结果消息),我们在我们网站的平台上报告了一组数字实验。该实验表明,我们的矩阵产品算法的执行时间比现有的算法少,而它也使用较少的资源。
This paper is focused on designing efficient parallel matrix-product algorithms for heterogeneous master-worker platforms. While matrix-product is well-understood for homogeneous 2D-arrays of processors (e.g., Cannon algorithm and ScaLAPACK outer product algorithm), there are three key hypotheses that render our work original and innovative: - Centralized data. We assume that all matrix files originate from, and must be returned to, the master. The master distributes data and computations to the workers while in ScaLAPACK, input and output matrices are supposed to be equally distributed among participating resources beforehand). Typically, our approach is useful in the context of speeding up MATLAB or SCILAB clients running on a server (which acts as the master and initial repository of files). - Heterogeneous star-shaped platforms. We target fully heterogeneous platforms, where computational resources have different computing powers. Also, the workers are connected to the master by links of different capacities. This framework is realistic when deploying the application from the server, which is responsible for enrolling authorized resources. - Limited memory. As we investigate the parallelization of large problems, we cannot assume that full matrix column blocks can be stored in the worker memories and be re-used for subsequent updates (as in ScaLAPACK). We have devised efficient algorithms for resource selection (deciding which workers to enroll) and communication ordering (both for input and result messages), and we report a set of numerical experiments on a platform at our site. The experiments show that our matrix-product algorithm has smaller execution times than existing ones, while it also uses fewer resources.