HERO: hessian-enhanced robust optimization for unifying and improving generalization and quantization performance

HERO: hessian-enhanced robust optimization for unifying and improving generalization and quantization performance
复制标题

HERO:hessian 增强的鲁棒优化,用于统一和提高泛化和量化性能

DOI:
10.1145/3489517.3530678
复制
发表时间:
2022
期刊:
The 59th ACM/IEEE Design Automation Conference
影响因子:
--
通讯作者:
Chen, Yiran
Chen, Yiran
中科院分区:
--
文献类型:
--
作者:
Yang, Huanrui;Yang, Xiaoxuan;Gong, Neil Zhenqiang;Chen, Yiran

文献摘要

参考文献

相似文献

随着最近在移动的和边缘设备上部署神经网络模型的需求,期望提高模型在看不见的测试数据上的泛化能力,以及增强模型在定点量化下的鲁棒性以进行有效部署。然而,最小化训练损失对泛化和量化性能几乎没有保证。在这项工作中,我们满足需要,同时提高泛化和量化性能的理论统一的框架下,提高模型的鲁棒性有界的权重扰动和最小化的Hessian矩阵的特征值相对于模型的权重。因此,我们提出了HERO,一种Hessian增强的鲁棒优化方法,通过基于梯度的训练过程来最小化Hessian特征值,同时提高泛化和量化性能。HERO使测试准确度提高了3.8%,在80%的训练标签扰动下准确度提高了30%,并且在广泛的精度范围内具有最佳的训练后量化准确度,包括在各种数据集上的通用模型架构中,比SGD训练模型的准确度提高了> 10%。
With the recent demand of deploying neural network models on mobile and edge devices, it is desired to improve the model's generalizability on unseen testing data, as well as enhance the model's robustness under fixed-point quantization for efficient deployment. Minimizing the training loss, however, provides few guarantees on the generalization and quantization performance. In this work, we fulfill the need of improving generalization and quantization performance simultaneously by theoretically unifying them under the framework of improving the model's robustness against bounded weight perturbation and minimizing the eigenvalues of the Hessian matrix with respect to model weights. We therefore propose HERO, a Hessian-enhanced robust optimization method, to minimize the Hessian eigenvalues through a gradient-based training process, simultaneously improving the generalization and quantization performance. HERO enables up to a 3.8% gain on test accuracy, up to 30% higher accuracy under 80% training label perturbation, and the best post-training quantization accuracy across a wide range of precision, including a > 10% accuracy improvement over SGD-trained models for common model architectures on various datasets.
DOI: --
发表时间: 2021-02
期刊: ArXiv
影响因子: --
作者:
Huanrui Yang;Lin Duan;Yiran Chen;Hai Li
通讯作者: Huanrui Yang;Lin Duan;Yiran Chen;Hai Li
利用深度学习和海运物流大数据预测航运市场状况的研究
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
和田 祐次郎;河原 大輝;濱田 邦裕
通讯作者: 濱田 邦裕
用于量化鲁棒性的梯度 ?1 正则化
DOI: --
发表时间: 2020
期刊: arXiv.org
影响因子: --
作者:
Milad Alizadeh;A. Behboodi;M. V. Baalen;Christos Louizos;Tijmen Blankevoort;M. Welling
通讯作者: M. Welling
能够使用深度学习从面部照片中区分阿尔茨海默病
DOI: --
发表时间: 2020
期刊:
影响因子: --
作者:
Manabu Kokubo;Akihiro Hirashiki;Takahiro Kamihara;Atsuya Shimizu;Hidenori Arai;亀山祐美,亀山征史,深澤誠,飯塚友道,飯島勝矢,田中友規,矢可部満隆,小島太郎,小川純人,秋下雅弘
通讯作者: 亀山祐美,亀山征史,深澤誠,飯塚友道,飯島勝矢,田中友規,矢可部満隆,小島太郎,小川純人,秋下雅弘