Betty: An Automatic Differentiation Library for Multilevel Optimization

Betty: An Automatic Differentiation Library for Multilevel Optimization
复制标题

DOI:
10.48550/arxiv.2207.02849
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Sang Keun Choe;W. Neiswanger;P. Xie;Eric P. Xing
Sang Keun Choe;W. Neiswanger;P. Xie;Eric P. Xing
中科院分区:
其他
文献类型:
--
作者:
Sang Keun Choe;W. Neiswanger;P. Xie;Eric P. Xing

文献摘要

相似文献

基于梯度的多级优化(MLO)作为研究众多问题的框架而受到关注,从超参数优化和元学习到神经架构搜索和强化学习。然而,MLO 中的梯度是通过链式法则组合最佳响应雅可比行列式获得的,其实现起来非常困难,并且需要大量内存/计算。我们为缩小这一差距迈出了第一步,引入了 Betty,一个用于大规模 MLO 的软件库。其核心是,我们为 MLO 设计了一种新颖的数据流图,它使我们能够 (1) 为 MLO 开发高效的自动微分,将计算复杂度从 O(d^3) 降低到 O(d^2),(2) 合并系统支持,例如混合精度和数据并行训练以实现可扩展性,(3) 促进任意复杂度的 MLO 程序的实现,同时允许为不同的算法和系统设计选择提供模块化接口。我们凭经验证明,Betty 可用于实现一系列 MLO 程序,同时在多个基准测试中,与现有实现相比,测试精度提高了 11%,GPU 内存使用量减少了 14%,训练时间减少了 20%。我们还展示了 Betty 能够将 MLO 扩展到具有数亿参数的模型。我们在 https://github.com/leopard-ai/betty 开源代码。
Gradient-based multilevel optimization (MLO) has gained attention as a framework for studying numerous problems, ranging from hyperparameter optimization and meta-learning to neural architecture search and reinforcement learning. However, gradients in MLO, which are obtained by composing best-response Jacobians via the chain rule, are notoriously difficult to implement and memory/compute intensive. We take an initial step towards closing this gap by introducing Betty, a software library for large-scale MLO. At its core, we devise a novel dataflow graph for MLO, which allows us to (1) develop efficient automatic differentiation for MLO that reduces the computational complexity from O(d^3) to O(d^2), (2) incorporate systems support such as mixed-precision and data-parallel training for scalability, and (3) facilitate implementation of MLO programs of arbitrary complexity while allowing a modular interface for diverse algorithmic and systems design choices. We empirically demonstrate that Betty can be used to implement an array of MLO programs, while also observing up to 11% increase in test accuracy, 14% decrease in GPU memory usage, and 20% decrease in training wall time over existing implementations on multiple benchmarks. We also showcase that Betty enables scaling MLO to models with hundreds of millions of parameters. We open-source the code at https://github.com/leopard-ai/betty.