Blessing of Class Diversity in Pre-training

Blessing of Class Diversity in Pre-training
复制标题

DOI:
10.48550/arxiv.2209.03447
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Yulai Zhao;Jianshu Chen;S. Du
Yulai Zhao;Jianshu Chen;S. Du
中科院分区:
其他
文献类型:
--
作者:
Yulai Zhao;Jianshu Chen;S. Du

文献摘要

相似文献

本文提出了一种新的统计分析,旨在解释最近的上级成就的预训练技术在自然语言处理(NLP)。我们证明,当预训练任务的类(例如,掩码语言模型任务中的不同单词)是足够多样的,在这个意义上,预训练中最后一个线性层的最小奇异值(表示为$\tilde{\nu}$)很大,那么预训练可以显著提高下游任务的样本效率。特别地,我们证明了转移学习的超额风险具有$O\left(\frac{1}{\nu} \sqrt{n}}\right)$率,而标准监督学习的超额风险率为$O\left(\frac{1}{\sqrt{m}}\right)$率。这里,$n$是预训练数据的数量,$m$是下游任务中的数据数量,通常为$n \gg m$。我们的证明依赖于一个向量形式的Rademacher复杂性链规则分解复合函数类和修改后的自一致性条件。这些技术可以是独立的兴趣。
This paper presents a new statistical analysis aiming to explain the recent superior achievements of the pre-training techniques in natural language processing (NLP). We prove that when the classes of the pre-training task (e.g., different words in the masked language model task) are sufficiently diverse, in the sense that the least singular value of the last linear layer in pre-training (denoted as $\tilde{\nu}$) is large, then pre-training can significantly improve the sample efficiency of downstream tasks. Specially, we show the transfer learning excess risk enjoys an $O\left(\frac{1}{\tilde{\nu} \sqrt{n}}\right)$ rate, in contrast to the $O\left(\frac{1}{\sqrt{m}}\right)$ rate in the standard supervised learning. Here, $n$ is the number of pre-training data and $m$ is the number of data in the downstream task, and typically $n \gg m$. Our proof relies on a vector-form Rademacher complexity chain rule for disassembling composite function classes and a modified self-concordance condition. These techniques can be of independent interest.