Performance-Aware Mutual Knowledge Distillation for Improving Neural Architecture Search

Performance-Aware Mutual Knowledge Distillation for Improving Neural Architecture Search
复制标题

DOI:
10.1109/cvpr52688.2022.01162
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
P. Xie;Xuefeng Du
P. Xie;Xuefeng Du
中科院分区:
其他
文献类型:
--
作者:
P. Xie;Xuefeng Du

文献摘要

相似文献

知识提取在改进神经体系结构搜索(NAS)方面显示出了巨大的有效性。互知识蒸馏(MKD)是一组模型相互产生知识相互训练的过程,在许多应用中都取得了很好的效果。在现有的MKD方法中,模型之间的相互知识提炼是在没有仔细检查的情况下进行的:允许表现较差的模型生成知识来训练表现较好的模型,这可能导致集体失败。为了解决这个问题,我们提出了一种面向NAS的性能感知MKD(PAMKD)方法,其中只有当$A$的性能好于B时,才允许模型$A产生的知识训练$B$。我们提出了一个三级优化框架来制定PAMKD,其中端到端执行三个学习阶段:1)每个模型独立地训练初始模型;2)在验证集上评估初始模型,而性能较好的模型生成知识来训练性能较差的模型;3)通过最小化验证损失来更新体系结构。在多种数据集上的实验结果表明,该方法是有效的。
Knowledge distillation has shown great effectiveness for improving neural architecture search (NAS). Mutual knowledge distillation (MKD), where a group of models mutually generate knowledge to train each other, has achieved promising results in many applications. In existing MKD methods, mutual knowledge distillation is performed between models without scrutiny: a worse-performing model is allowed to generate knowledge to train a better-performing model, which may lead to collective failures. To address this problem, we propose a performance-aware MKD (PAMKD) approach for NAS, where knowledge generated by model $A$ is allowed to train model $B$ only if the performance of $A$ is better than B. We propose a three-level optimization framework to formulate PAMKD, where three learning stages are performed end-to-end: 1) each model trains an initial model independently; 2) the initial models are evaluated on a validation set and better-performing models generate knowledge to train worse-performing models; 3) architectures are updated by minimizing a validation loss. Experimental results on a variety of datasets demonstrate that our method is effective.