Characterizing and Understanding the Generalization Error of Transfer Learning with Gibbs Algorithm

Characterizing and Understanding the Generalization Error of Transfer Learning with Gibbs Algorithm
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuheng Bu;Gholamali Aminian;L. Toni;Miguel L. Rodrigues;G. Wornell
Yuheng Bu;Gholamali Aminian;L. Toni;Miguel L. Rodrigues;G. Wornell
中科院分区:
其他
文献类型:
--
作者:
Yuheng Bu;Gholamali Aminian;L. Toni;Miguel L. Rodrigues;G. Wornell

文献摘要

相似文献

我们通过专注于两种流行的转移学习方法,分别$ \ alpha $ - 加权 - erm和两阶段的erm,对基于吉布斯的转移学习算法的概括能力提供了信息理论分析。我们的关键结果是使用条件对称的KL信息在输出假设和给定源样本的目标训练样本之间对概括行为进行了精确表征。我们的结果还可以应用于上述两个Gibbs算法上的新型无分布概括上限。我们的方法是多才多艺的,因为它也表征了渐近制度中这两种Gibbs算法的概括错误和多余的风险,在那里它们分别融合到$ \ alpha $ - 加权词和两个阶段。基于我们的理论结果,我们表明,转移学习的好处可以视为偏见 - 差异权衡,而偏见是由源分布和缺乏目标样本引起的差异引起的。我们认为,这种观点可以指导在实践中选择转移学习算法的选择。
We provide an information-theoretic analysis of the generalization ability of Gibbs-based transfer learning algorithms by focusing on two popular transfer learning approaches, $\alpha$-weighted-ERM and two-stage-ERM. Our key result is an exact characterization of the generalization behaviour using the conditional symmetrized KL information between the output hypothesis and the target training samples given the source samples. Our results can also be applied to provide novel distribution-free generalization error upper bounds on these two aforementioned Gibbs algorithms. Our approach is versatile, as it also characterizes the generalization errors and excess risks of these two Gibbs algorithms in the asymptotic regime, where they converge to the $\alpha$-weighted-ERM and two-stage-ERM, respectively. Based on our theoretical results, we show that the benefits of transfer learning can be viewed as a bias-variance trade-off, with the bias induced by the source distribution and the variance induced by the lack of target samples. We believe this viewpoint can guide the choice of transfer learning algorithms in practice.