Parallel Software for Million-scale Exact Kernel Regression
Parallel Software for Million-scale Exact Kernel Regression
复制标题
DOI:
10.1145/3577193.3593737
复制
发表时间:
2023-06
期刊:
影响因子:
--
通讯作者:
Yu Chen;Lucca Skon;James R. McCombs;Zhenming Liu;A. Stathopoulos
中科院分区:
文献类型:
--
作者:
Yu Chen;Lucca Skon;James R. McCombs;Zhenming Liu;A. Stathopoulos
We present the design and the implementation of a kernel principal component regression software that handles training datasets with a million or more observations. Kernel regressions are nonlinear and interpretable models that have wide downstream applications, and are shown to have a close connection to deep learning. Nevertheless, the exact regression of large-scale kernel models using currently available software has been notoriously difficult because it is both compute and memory intensive and it requires extensive tuning of hyperparameters. While in computational science distributed computing and iterative methods have been a mainstay of large scale software, they have not been widely adopted in kernel learning. Our software leverages existing high performance computing (HPC) techniques and develops new ones that address cross-cutting constraints between HPC and learning algorithms. It integrates three major components: (a) a state-of-the-art parallel eigenvalue iterative solver, (b) a block matrix-vector multiplication routine that employs both multi-threading and distributed memory parallelism and can be performed on-the-fly under limited memory, and (c) a software pipeline consisting of Python front-ends that control the HPC backbone and the hyperparameter optimization through a boosting optimizer. We perform feasibility studies by running the entire ImageNet dataset and a large asset pricing dataset.