Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates

Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates
复制标题

DOI:
10.1145/3600006.3613152
复制
发表时间:
2023-09
期刊:
Proceedings of the 29th Symposium on Operating Systems Principles
影响因子:
--
通讯作者:
Insu Jang;Zhenning Yang;Zhen Zhang;Xin Jin;Mosharaf Chowdhury
Insu Jang;Zhenning Yang;Zhen Zhang;Xin Jin;Mosharaf Chowdhury
中科院分区:
其他
文献类型:
--
作者:
Insu Jang;Zhenning Yang;Zhen Zhang;Xin Jin;Mosharaf Chowdhury

文献摘要

被引文献

相似文献

Oobleck支持大型DNN模型的弹性分布式训练,并保证容错。它采用规划-执行协同设计方法,首先生成一组异构管道模板,并实例化至少f + 1个逻辑上等效的管道副本,以容忍任何f个同时发生的故障。在执行过程中,它依赖于已经跨副本复制的模型状态来提供快速恢复。Oobleck可证明地保证了最初创建的管道模板的某种组合可以用于在f个或更少的同时故障之后覆盖所有可用资源,从而避免了任何时候的资源空闲。对具有数十亿参数的大型DNN模型的评估表明,Oobleck提供了始终如一的高吞吐量,并且比Bamboo和Varuna等最先进的容错解决方案性能高出13.9倍。
Oobleck enables resilient distributed training of large DNN models with guaranteed fault tolerance. It takes a planning-execution co-design approach, where it first generates a set of heterogeneous pipeline templates and instantiates at least f + 1 logically equivalent pipeline replicas to tolerate any f simultaneous failures. During execution, it relies on already-replicated model states across the replicas to provide fast recovery. Oobleck provably guarantees that some combination of the initially created pipeline templates can be used to cover all available resources after f or fewer simultaneous failures, thereby avoiding resource idling at all times. Evaluation on large DNN models with billions of parameters shows that Oobleck provides consistently high throughput, and it outperforms state-of-the-art fault tolerance solutions like Bamboo and Varuna by up to 13.9×.