Robust fine-tuning of zero-shot models

Robust fine-tuning of zero-shot models
复制标题

DOI:
10.1109/cvpr52688.2022.00780
复制
发表时间:
2021-09
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Mitchell Wortsman;Gabriel Ilharco;Mike Li;Jong Wook Kim;Hannaneh Hajishirzi;Ali Farhadi;Hongseok Namkoong-H
Mitchell Wortsman;Gabriel Ilharco;Mike Li;Jong Wook Kim;Hannaneh Hajishirzi;Ali Farhadi;Hongseok Namkoong-H
中科院分区:
其他
文献类型:
--
作者:
Mitchell Wortsman;Gabriel Ilharco;Mike Li;Jong Wook Kim;Hannaneh Hajishirzi;Ali Farhadi;Hongseok Namkoong-H

文献摘要

被引文献

相似文献

大型预训练模型(如CLIP或ALIGN)在执行零次推理(即,而不对特定数据集进行微调)。虽然现有的微调方法大大提高了给定目标分布的准确性,但它们通常会降低对分布偏移的鲁棒性。我们通过引入一种简单而有效的方法来解决这种紧张局势,以提高鲁棒性,同时微调:集成的零拍摄和微调模型(WiSE-FT)的权重。与标准微调相比,WiSE-FT在分布偏移的情况下提供了较大的精度改进,同时保持了目标分布的高精度。在ImageNet和五个衍生的分布偏移上,WiSE-FT将分布偏移下的准确性提高了4到6个百分点(pp),同时将ImageNet的准确性提高了1.6个pp。WiSE-FT在一组不同的六个进一步的分布变化上实现了类似的大鲁棒性增益(2到23 pp),与常用的迁移学习数据集上的标准微调相比,准确性增益为0.8到3.3 pp。这些改进在微调或推断期间没有额外的计算成本。
Large pre-trained models such as CLIP or ALIGN offer consistent accuracy across a range of data distributions when performing zero-shot inference (i.e., without fine-tuning on a specific dataset). Although existing fine-tuning methods substantially improve accuracy on a given target distribution, they often reduce robustness to distribution shifts. We address this tension by introducing a simple and effective method for improving robustness while fine-tuning: ensembling the weights of the zero-shot and fine-tuned models (WiSE-FT). Compared to standard fine-tuning, WiSE-FT provides large accuracy improvements under distribution shift, while preserving high accuracy on the target distribution. On ImageNet and five derived distribution shifts, WiSE-FT improves accuracy under distribution shift by 4 to 6 percentage points (pp) over prior work while increasing ImageNet accuracy by 1.6 pp. WiSE-FT achieves similarly large robustness gains (2 to 23 pp) on a diverse set of six further distribution shifts, and accuracy gains of 0.8 to 3.3 pp compared to standard fine-tuning on commonly used transfer learning datasets. These improvements come at no additional computational cost during fine-tuning or inference.