Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear Algebra

Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear Algebra
复制标题

DOI:
10.1145/3448016.3457317
复制
发表时间:
2021-06
期刊:
Proceedings of the 2021 International Conference on Management of Data
影响因子:
--
通讯作者:
Shangyu Luo;Dimitrije Jankov;Binhang Yuan;C. Jermaine
Shangyu Luo;Dimitrije Jankov;Binhang Yuan;C. Jermaine
中科院分区:
其他
文献类型:
--
作者:
Shangyu Luo;Dimitrije Jankov;Binhang Yuan;C. Jermaine

文献摘要

相似文献

机器学习(ML)计算通常使用向量,矩阵或高维张量表示。这样的数据结构可以具有许多不同的实现,尤其是在分布式环境中:矩阵可以存储为行或列向量,不同尺寸或相关的瓷砖作为一组(rowindex,colindex,value)三倍。许多其他存储格式是可能的。格式的选择可以对ML计算的性能产生深远的影响。在本文中,我们提出了一个框架,以自动优化分布式环境中复杂的ML或线性代数(LA)计算的物理实现,开发用于解决此问题的算法,并通过分布式关系的原型显示,并通过分布式关系的原型显示数据库系统,我们的想法可以从根本上加快常见的ML和LA计算加快。
Machine learning (ML) computations are often expressed using vectors, matrices, or higher-dimensional tensors. Such data structures can have many different implementations, especially in a distributed environment: a matrix could be stored as row or column vectors, tiles of different sizes, or relationally, as a set of (rowIndex, colIndex, value) triples. Many other storage formats are possible. The choice of format can have a profound impact on the performance of a ML computation. In this paper, we propose a framework for automatic optimization of the physical implementation of a complex ML or linear algebra (LA) computation in a distributed environment, develop algorithms for solving this problem, and show, through a prototype on top of a distributed relational database system, that our ideas can radically speed up common ML and LA computations.