Data-Efficient Double-Win Lottery Tickets from Robust Pre-training

Data-Efficient Double-Win Lottery Tickets from Robust Pre-training
复制标题

DOI:
10.48550/arxiv.2206.04762
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Tianlong Chen;Zhenyu (Allen) Zhang;Sijia Liu;Yang Zhang;Shiyu Chang;Zhangyang Wang
Tianlong Chen;Zhenyu (Allen) Zhang;Sijia Liu;Yang Zhang;Shiyu Chang;Zhangyang Wang
中科院分区:
其他
文献类型:
--
作者:
Tianlong Chen;Zhenyu (Allen) Zhang;Sijia Liu;Yang Zhang;Shiyu Chang;Zhangyang Wang

文献摘要

相似文献

预训练是迁移学习在各种下游任务中广泛采用的起点。最近对彩票假设(LTH)的研究表明,这种巨大的预训练模型可以用极其稀疏的子网络(又名匹配子网络)代替,而不会牺牲可转移性。然而,实际的安全关键应用通常会提出标准传输之外的更具挑战性的要求,这也要求这些子网克服对抗性漏洞。在本文中,我们提出了一个更严格的概念,双赢彩票,其中预训练模型的定位子网络可以独立地转移到不同的下游任务上,在标准和对抗训练制度下,达到相同的标准和鲁棒泛化,就像完整的预训练模型可以做到的那样。我们全面研究了各种预训练机制,发现稳健的预训练倾向于制作更稀疏的双赢彩票,其性能优于标准对手。例如,在下游CIFAR-10/100数据集上,我们通过ImageNet的标准、快速对抗和对抗预训练识别出双赢匹配子网,稀疏度分别为89.26%/73.79%、89.26%/79.03%和91.41%/83.22%。此外,我们观察到,在实际数据有限(例如1%和10%)的下游方案下,获得的双赢彩票可以更有效地传输数据。我们的研究结果表明,鲁棒预训练的好处被彩票方案和数据有限的传输设置放大了。代码可在https://github.com/VITA-Group/Double-Win-LTH上获得。
Pre-training serves as a broadly adopted starting point for transfer learning on various downstream tasks. Recent investigations of lottery tickets hypothesis (LTH) demonstrate such enormous pre-trained models can be replaced by extremely sparse subnetworks (a.k.a. matching subnetworks) without sacrificing transferability. However, practical security-crucial applications usually pose more challenging requirements beyond standard transfer, which also demand these subnetworks to overcome adversarial vulnerability. In this paper, we formulate a more rigorous concept, Double-Win Lottery Tickets, in which a located subnetwork from a pre-trained model can be independently transferred on diverse downstream tasks, to reach BOTH the same standard and robust generalization, under BOTH standard and adversarial training regimes, as the full pre-trained model can do. We comprehensively examine various pre-training mechanisms and find that robust pre-training tends to craft sparser double-win lottery tickets with superior performance over the standard counterparts. For example, on downstream CIFAR-10/100 datasets, we identify double-win matching subnetworks with the standard, fast adversarial, and adversarial pre-training from ImageNet, at 89.26%/73.79%, 89.26%/79.03%, and 91.41%/83.22% sparsity, respectively. Furthermore, we observe the obtained double-win lottery tickets can be more data-efficient to transfer, under practical data-limited (e.g., 1% and 10%) downstream schemes. Our results show that the benefits from robust pre-training are amplified by the lottery ticket scheme, as well as the data-limited transfer setting. Codes are available at https://github.com/VITA-Group/Double-Win-LTH.