CAREER: Recursive Distributed Matrix and Tensor Decompositions on Neural Engines
CAREER: Recursive Distributed Matrix and Tensor Decompositions on Neural Engines
批准号:
2146509
负责人:
Panruo Wu
金额:
$52.87万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-03-01 至 2027-02-28
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Matrix and tensor decompositions are one of the most important building blocks for scientific computing and are increasingly important in data-centric computing and machine-learning models. The lack of software and algorithms that can efficiently deal with large data sets and exploit the ubiquitous availability of neural engines is holding back progress. Legacy distributed matrix packages based on complex data distribution schemes not only add friction in adoption in new areas but also impede the exploration of cutting-edge algorithms at scale. New exciting algorithms such as randomized linear algebra, structured matrix computation, and advanced eigen decompositions that are synergistic to neural engines remain unexplored, ad-hoc, or hard to use by non-experts in numerical analysis. New powerful architectures -- neural engines -- promise orders of magnitudes o performance and energy benefits but remain a challenge to use outside of neural networks. This proposal aims to create a unified software system to achieve high-performance, scalable, distributed matrix and tensor decompositions on neural engines through concerted research and development.This project addresses three research thrusts to achieve its goals. A) In contrast to conventional arithmetic-centric algorithm design, this research focuses on communication-efficient algorithm variants. A central challenge in realizing the proposed goals is the avoidance, and management, of data movement. Computation speed has become amazingly fast on neural engines, while data movement latency and bandwidth lag far behind and the gap is widening. B) Incorporation of neural engines to state-of-the-art numerical algorithms. Recent numerical analysis has seen some exciting developments in randomized algorithms, low-precision direct decomposition as a preconditioner, and novel polar decomposition-based spectral divide-and-conquer methods for eigensystems. These new developments are not only exciting by themselves, but they have the potential to exploit neural engines especially well and blend with communication-centric algorithms naturally. C) Exploration of Universal Distributed Array (UDA), a new data structure based on a multi-dimensional cyclic data distribution scheme, to achieve load balancing, scalability, and unified support for all matrix and tensor decompositions. This proposal extends the cyclic data-distribution scheme to support communication-efficient algorithms including recursive algorithms due to flexible alignment, and to multi-dimensional to support tensor decomposition. The project will develop efficient, scalable, and easy-to-use communication and computational primitives on distributed neural engines and will include the most useful matrix/tensor decomposition algorithms as a composable and extensible library.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Fast Symmetric Eigenvalue Decomposition via WY Representation on Tensor Core
通过张量核心上的 WY 表示进行快速对称特征值分解
DOI:
10.1145/3572848.3577516
发表时间:
2023
期刊:
ACM
影响因子:
--
作者:
[Zhang, Shaoshuai, Shah, Ruchi, Ootomo, Hiroyuki, Yokota, Rio, Wu, Panruo]
通讯作者:
Wu, Panruo
海外基金