Learning to Minimize the Remainder in Supervised Learning

Learning to Minimize the Remainder in Supervised Learning
复制标题

DOI:
10.1109/tmm.2022.3158066
复制
发表时间:
2022-01
影响因子:
7.3
通讯作者:
Yan Luo;Yongkang Wong;Mohan S. Kankanhalli;Qi Zhao-
Yan Luo;Yongkang Wong;Mohan S. Kankanhalli;Qi Zhao-
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yan Luo;Yongkang Wong;Mohan S. Kankanhalli;Qi Zhao-

文献摘要

相似文献

深度学习方法的学习过程通常会在多次迭代中更新模型参数。每一次迭代都可以看作泰勒级数展开的一阶近似。为简单起见,在学习过程中通常忽略由高阶项组成的余项。这种学习方案支持各种基于多媒体的应用,如图像检索、推荐系统和视频搜索。通常,多媒体数据(例如图像)是语义丰富的高维数据,因此近似的剩余部分可能是非零的。在这项工作中,我们认为其余的是信息的,并研究它如何影响学习过程。为此,我们提出了一种新的学习方法,即梯度调整学习(GAL),利用从过去训练迭代中学习到的知识来调整香草梯度,从而最小化剩余项并改进逼近。所提出的GAL与模型和优化器无关,并且很容易适应标准的学习框架。使用最先进的模型和优化器,从图像分类、目标检测和回归三个方面对该算法进行了评估。实验表明,所提出的GAL一致地改进了被评估的模型,而烧蚀研究则验证了所提出的GAL的各个方面。代码可在https://github.com/luoyan407/gradient_adjustment.git.上获得
The learning process of deep learning methods usually updates the model’s parameters in multiple iterations. Each iteration can be viewed as the first-order approximation of Taylor’s series expansion. The remainder, which consists of higher-order terms, is usually ignored in the learning process for simplicity. This learning scheme empowers various multimedia-based applications, such as image retrieval, recommendation system, and video search. Generally, multimedia data (e.g. images) are semantics-rich and high-dimensional, hence the remainders of approximations are possibly non-zero. In this work, we consider that the remainder is informative and study how it affects the learning process. To this end, we propose a new learning approach, namely gradient adjustment learning (GAL), to leverage the knowledge learned from the past training iterations to adjust vanilla gradients, such that the remainders are minimized and the approximations are improved. The proposed GAL is model- and optimizer-agnostic, and is easy to adapt to the standard learning framework. It is evaluated on three tasks, i.e. image classification, object detection, and regression, with state-of-the-art models and optimizers. The experiments show that the proposed GAL consistently enhances the evaluated models, whereas the ablation studies validate various aspects of the proposed GAL. The code is available at https://github.com/luoyan407/gradient_adjustment.git.