Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning

Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning
复制标题

DOI:
10.14778/3476249.3476284
复制
发表时间:
2021-07
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Side Li;Arun Kumar
Side Li;Arun Kumar
中科院分区:
其他
文献类型:
--
作者:
Side Li;Arun Kumar

文献摘要

相似文献

许多使用大规模机器学习(ML)的应用越来越倾向于为子组使用不同的模型(例如,国家)以提高准确性、公平性或其他必要条件。我们将这种新兴的流行实践称为组学习,类似于SQL中的GROUP BY,尽管是用于ML训练而不是SQL聚合。从系统的角度来看,这种做法加剧了ML模型选择的数据密集型工作量(例如,超参数调整)。通常,可能需要训练数千个模型,需要高吞吐量并行执行。遗憾的是,今天的大多数ML系统都专注于一次训练一个模型,或者最好是并行化超参数调优。这种现状导致资源浪费、低吞吐量和高运行时间。在这项工作中,我们从数据系统的角度出发,为三种流行的ML类别(线性模型,神经网络和梯度提升决策树)实现和优化组学习迈出了第一步。通过分析和经验,我们比较了当今执行此工作负载的标准方法:任务并行和数据并行。我们发现两者都不是普遍占主导地位的。我们提出了一种新的混合方法,我们称之为分组学习,它使用一种新形式的并行梯度下降来避免通信和I/O中的冗余,我们称之为梯度累积算法(GAP)。我们将我们的想法原型化为一个系统,我们称之为Kingpin,该系统构建在现有ML工具和灵活的大规模并行运行时Ray之上。对大型ML基准数据集的广泛经验评估表明,Kingpin匹配或比最先进的ML系统快4倍到14倍,包括Ray的本地执行和PyTorch DDP。
Many applications that use large-scale machine learning (ML) increasingly prefer different models for subgroups (e.g., countries) to improve accuracy, fairness, or other desiderata. We call this emerging popular practice learning over groups , analogizing to GROUP BY in SQL, albeit for ML training instead of SQL aggregates. From the systems standpoint, this practice compounds the already data-intensive workload of ML model selection (e.g., hyperparameter tuning). Often, thousands of models may need to be trained, necessitating high-throughput parallel execution. Alas, most ML systems today focus on training one model at a time or at best, parallelizing hyperparameter tuning. This status quo leads to resource wastage, low throughput, and high runtimes. In this work, we take the first step towards enabling and optimizing learning over groups from the data systems standpoint for three popular classes of ML: linear models, neural networks, and gradient-boosted decision trees. Analytically and empirically, we compare standard approaches to execute this workload today: task-parallelism and data-parallelism. We find neither is universally dominant. We put forth a novel hybrid approach we call grouped learning that avoids redundancy in communications and I/O using a novel form of parallel gradient descent we call Gradient Accumulation Parallelism (GAP). We prototype our ideas into a system we call Kingpin built on top of existing ML tools and the flexible massively-parallel runtime Ray. An extensive empirical evaluation on large ML benchmark datasets shows that Kingpin matches or is 4x to 14x faster than state-of-the-art ML systems, including Ray's native execution and PyTorch DDP.