DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
复制标题

DOI:
--
复制
发表时间:
2022-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Samyam Rajbhandari;Conglong Li;Z. Yao;Minjia Zhang;Reza Yazdani Aminabadi;A. A. Awan-A.;Jeff Rasley;Yuxiong He
Samyam Rajbhandari;Conglong Li;Z. Yao;Minjia Zhang;Reza Yazdani Aminabadi;A. A. Awan-A.;Jeff Rasley;Yuxiong He
中科院分区:
其他
文献类型:
--
作者:
Samyam Rajbhandari;Conglong Li;Z. Yao;Minjia Zhang;Reza Yazdani Aminabadi;A. A. Awan-A.;Jeff Rasley;Yuxiong He

文献摘要

被引文献

相似文献

随着大型密集模型的训练达到硬件资源可用性和能力的极限,与质量等效的密集模型相比,混合专家(MoE)模型显著降低了训练成本,成为最有前途的模型体系结构之一。从编码器-解码器模型(以前的工作)到自动攻击语言模型(与并行探索一起工作)的5倍节省的训练成本被证明。然而,由于模型尺寸更大,结构独特,如何提供快速的MoE模型推理仍然是一个挑战和未解决的问题,限制了其实际应用。为了解决这个问题,我们提出了DeepSpeed-MoE,这是一个端到端的MoE训练和推理解决方案,作为DeepSpeed库的一部分,包括新颖的MoE架构设计和模型压缩技术,可将MoE模型大小减少3.7倍,以及一个高度优化的推理系统,与现有的MoE推理解决方案相比,它提供了7.3倍的延迟和成本。与质量相当的密集模型相比,DeepSpeed-MoE提供了前所未有的规模和效率,为大规模MoE模型提供了高达4.5倍的速度和9倍的成本推断。我们希望我们的创新和系统有助于在大型模型领域开辟一条有前途的新方向,从密集到稀疏的MoE模型的转变,在这种情况下,用更少的资源训练和部署更高质量的模型变得更加可能。
As the training of giant dense models hits the boundary on the availability and capability of the hardware resources today, Mixture-of-Experts (MoE) models become one of the most promising model architectures due to their significant training cost reduction compared to a quality-equivalent dense model. Its training cost saving is demonstrated from encoder-decoder models (prior works) to a 5x saving for auto-aggressive language models (this work along with parallel explorations). However, due to the much larger model size and unique architecture, how to provide fast MoE model inference remains challenging and unsolved, limiting its practical usage. To tackle this, we present DeepSpeed-MoE, an end-to-end MoE training and inference solution as part of the DeepSpeed library, including novel MoE architecture designs and model compression techniques that reduce MoE model size by up to 3.7x, and a highly optimized inference system that provides 7.3x better latency and cost compared to existing MoE inference solutions. DeepSpeed-MoE offers an unprecedented scale and efficiency to serve massive MoE models with up to 4.5x faster and 9x cheaper inference compared to quality-equivalent dense models. We hope our innovations and systems help open a promising path to new directions in the large model landscape, a shift from dense to sparse MoE models, where training and deploying higher-quality models with fewer resources becomes more widely possible.