Continual Learning of Generative Models with Limited Data: From Wasserstein-1 Barycenter to Adaptive Coalescence

Continual Learning of Generative Models with Limited Data: From Wasserstein-1 Barycenter to Adaptive Coalescence
复制标题

DOI:
10.1109/tnnls.2023.3251096
复制
发表时间:
2021-01
影响因子:
10.4
通讯作者:
M. Dedeoglu;Sen Lin;Zhaofeng Zhang;Junshan Zhang
M. Dedeoglu;Sen Lin;Zhaofeng Zhang;Junshan Zhang
中科院分区:
计算机科学1区
文献类型:
--
作者:
M. Dedeoglu;Sen Lin;Zhaofeng Zhang;Junshan Zhang

文献摘要

相似文献

对于数据和计算能力有限的网络边缘节点来说,学习生成模型具有挑战性。由于相似环境中的任务具有模型相似性,因此利用来自其他边缘节点的预训练生成模型是合理的。本研究呼吁针对Wasserstein-1生成对抗网络(WGAN)量身定制的最优传输理论,旨在开发一个框架,该框架使用边缘节点的本地数据系统地优化生成模型的持续学习,同时利用预训练生成模型的自适应合并。具体来说,通过将来自其他节点的知识转移视为以其预训练模型为中心的Wasserstein球,生成模型的持续学习被视为约束优化问题,该问题进一步简化为Wasserstein-1重心问题。因此,制定了一个两阶段的办法:1)离线计算预训练模型之间的重心,其中位移插值被用作经由“递归”WGAN配置找到自适应重心的理论基础,以及2)离线计算的重心被用作持续学习的元模型初始化,然后,使用目标边缘节点处的局部样本执行快速自适应以找到生成模型。最后,提出了一种基于权值和量化阈值联合优化的权值三值化方法,进一步压缩生成模型。大量的实验研究证实了所提出的框架的有效性。
Learning generative models is challenging for a network edge node with limited data and computing power. Since tasks in similar environments share a model similarity, it is plausible to leverage pretrained generative models from other edge nodes. Appealing to optimal transport theory tailored toward Wasserstein-1 generative adversarial networks (WGANs), this study aims to develop a framework that systematically optimizes continual learning of generative models using local data at the edge node while exploiting adaptive coalescence of pretrained generative models. Specifically, by treating the knowledge transfer from other nodes as Wasserstein balls centered around their pretrained models, continual learning of generative models is cast as a constrained optimization problem, which is further reduced to a Wasserstein-1 barycenter problem. A two-stage approach is devised accordingly: 1) the barycenters among the pretrained models are computed offline, where displacement interpolation is used as the theoretic foundation for finding adaptive barycenters via a "recursive" WGAN configuration and 2) the barycenter computed offline is used as metamodel initialization for continual learning, and then, fast adaptation is carried out to find the generative model using the local samples at the target edge node. Finally, a weight ternarization method, based on joint optimization of weights and threshold for quantization, is developed to compress the generative model further. Extensive experimental studies corroborate the effectiveness of the proposed framework.