LambdaNetworks: Modeling Long-Range Interactions Without Attention

LambdaNetworks: Modeling Long-Range Interactions Without Attention
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Irwan Bello
Irwan Bello
中科院分区:
其他
文献类型:
--
作者:
Irwan Bello

文献摘要

被引文献

相似文献

我们提出了lambda层--自我注意力的替代框架--用于捕获输入和结构化上下文信息(例如,被其他像素包围的像素)之间的远程交互。Lambda层通过将可用上下文转换为线性函数(称为Lambda)并将这些线性函数分别应用于每个输入来捕获此类交互。与线性注意力类似,lambda层绕过了昂贵的注意力地图,但相比之下,它们对基于内容和位置的交互进行了建模,这使得它们能够应用于大型结构化输入,如图像。由此产生的神经网络架构LambdaNetworks在ImageNet分类、COCO对象检测和COCO实例分割方面明显优于其卷积和注意力对手,同时计算效率更高。此外,我们还设计了LambdaResNets,这是一系列跨不同规模的混合架构,大大提高了图像分类模型的速度-准确性权衡。LambdaResNets在ImageNet上达到了出色的精度,同时比现代机器学习加速器上流行的EfficientNets快3.2 - 4.4倍。当使用额外的1.3亿个伪标记图像进行训练时,LambdaResNets的速度比相应的EfficientNet检查点提高了9.5倍。
We present lambda layers -- an alternative framework to self-attention -- for capturing long-range interactions between an input and structured contextual information (e.g. a pixel surrounded by other pixels). Lambda layers capture such interactions by transforming available contexts into linear functions, termed lambdas, and applying these linear functions to each input separately. Similar to linear attention, lambda layers bypass expensive attention maps, but in contrast, they model both content and position-based interactions which enables their application to large structured inputs such as images. The resulting neural network architectures, LambdaNetworks, significantly outperform their convolutional and attentional counterparts on ImageNet classification, COCO object detection and COCO instance segmentation, while being more computationally efficient. Additionally, we design LambdaResNets, a family of hybrid architectures across different scales, that considerably improves the speed-accuracy tradeoff of image classification models. LambdaResNets reach excellent accuracies on ImageNet while being 3.2 - 4.4x faster than the popular EfficientNets on modern machine learning accelerators. When training with an additional 130M pseudo-labeled images, LambdaResNets achieve up to a 9.5x speed-up over the corresponding EfficientNet checkpoints.