BASE Layers: Simplifying Training of Large, Sparse Models

BASE Layers: Simplifying Training of Large, Sparse Models
复制标题

DOI:
--
复制
发表时间:
2021-03
期刊:
--
影响因子:
--
通讯作者:
M. Lewis;Shruti Bhosale;Tim Dettmers;Naman Goyal;Luke Zettlemoyer
M. Lewis;Shruti Bhosale;Tim Dettmers;Naman Goyal;Luke Zettlemoyer
中科院分区:
其他
文献类型:
--
作者:
M. Lewis;Shruti Bhosale;Tim Dettmers;Naman Goyal;Luke Zettlemoyer

文献摘要

被引文献

相似文献

我们为大型语言模型引入了一个新的专家平衡分配(BASE)层,大大简化了现有的高容量稀疏层。稀疏层可以通过将每个令牌路由到仅包含一小部分模型参数的专用专家模块来显着提高训练和推理的效率。然而,它可以是难以学习的平衡路由功能,充分利用可用的专家,现有的方法通常使用路由算法或辅助专家平衡损失函数。相比之下,我们制定令牌专家分配作为一个线性分配问题,允许最佳分配,其中每个专家收到相同数量的令牌。这种最优分配方案通过保证平衡的计算负载来提高效率,并且通过不需要任何新的超参数或辅助损失来简化训练。代码在https://github.com/pytorch/fairseq/上公开发布
We introduce a new balanced assignment of experts (BASE) layer for large language models that greatly simplifies existing high capacity sparse layers. Sparse layers can dramatically improve the efficiency of training and inference by routing each token to specialized expert modules that contain only a small fraction of the model parameters. However, it can be difficult to learn balanced routing functions that make full use of the available experts; existing approaches typically use routing heuristics or auxiliary expert-balancing loss functions. In contrast, we formulate token-to-expert allocation as a linear assignment problem, allowing an optimal assignment in which each expert receives an equal number of tokens. This optimal assignment scheme improves efficiency by guaranteeing balanced compute loads, and also simplifies training by not requiring any new hyperparameters or auxiliary losses. Code is publicly released at https://github.com/pytorch/fairseq/