Representation Learning Beyond Linear Prediction Functions

Representation Learning Beyond Linear Prediction Functions
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
2023 International Conference on Integrated Intelligence and Communication Systems (ICIICS)
影响因子:
--
通讯作者:
Ziping Xu;Ambuj Tewari
Ziping Xu;Ambuj Tewari
中科院分区:
其他
文献类型:
--
作者:
Ziping Xu;Ambuj Tewari

文献摘要

被引文献

相似文献

关于表示理论学习理论的最新论文表明,从一组源任务概括到目标任务时,数量的重要性。这些论文中的大多数都认为,对于源和目标任务,函数映射共享表示为预测是线性的。在实践中,深度学习的研究人员根据新任务的难度使用了仔细的模型,使用了不同数量的额外层。这激发了我们询问当源任务和目标任务使用不同的预测功能空间以外的线性功能时,是否可以实现多样性。我们表明,即使目标任务使用具有多层的神经网络,只要源任务使用线性功能,多样性也会存在。如果源任务使用非线性预测函数,我们通过表明具有Relu激活功能的Depth-1神经网络指数上需要许多源任务来实现多样性来提供负面结果。对于通用功能类别,我们发现Eluder维度对多样性所需的任务数量进行了下限。我们的理论结果意味着更简单的任务可以更好地推广。尽管我们的理论结果是针对经验风险的全球最小化器显示的,但它们的定性预测仍然是基于梯度的优化算法的正确预测,这是我们对深神经网络的模拟验证的。
Recent papers on the theory of representation learning has shown the importance of a quantity called diversity when generalizing from a set of source tasks to a target task. Most of these papers assume that the function mapping shared representations to predictions is linear, for both source and target tasks. In practice, researchers in deep learning use different numbers of extra layers following the pretrained model based on the difficulty of the new task. This motivates us to ask whether diversity can be achieved when source tasks and the target task use different prediction function spaces beyond linear functions. We show that diversity holds even if the target task uses a neural network with multiple layers, as long as source tasks use linear functions. If source tasks use nonlinear prediction functions, we provide a negative result by showing that depth-1 neural networks with ReLu activation function need exponentially many source tasks to achieve diversity. For a general function class, we find that eluder dimension gives a lower bound on the number of tasks required for diversity. Our theoretical results imply that simpler tasks generalize better. Though our theoretical results are shown for the global minimizer of empirical risks, their qualitative predictions still hold true for gradient-based optimization algorithms as verified by our simulations on deep neural networks.