课题基金 / 基金详情

Robust Preconditioned Gradient Descent Algorithms for Deep Learning

Robust Preconditioned Gradient Descent Algorithms for Deep Learning
用于深度学习的鲁棒预条件梯度下降算法
批准号:
2208314
负责人:
Qiang Ye
金额:
$33.6万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2025-07-31

项目摘要

项目成果

Qiang Ye的其他基金

相似基金

相关文献

中文摘要
翻译
深度学习是人工智能和机器学习研究的前沿,影响着计算机视觉、语音识别、自然语言处理和生物信息学等数据科学的各种应用。深度神经网络学习中的一个关键挑战是模型优化,模型优化用于网络训练。然而,传统的优化算法并不适用,这主要是由于深度神经网络的高度复杂性和非线性。本项目的目标是开发新的稳健优化算法,能够有效地解决这些困难,并在实践中更有效地训练深度学习模型。该项目还涉及将这项工作应用于药物设计中使用的等效化学表示的翻译以及用于不确定性量化的贝叶斯推理。作为该项目的一部分,研究生和本科生将接受深度学习研究方面的培训,软件将被开发并免费提供。该项目包括开发两类新的优化算法,它们建立在传统的预条件和共轭梯度方法的框架上,但吸收了一些成功的专业深度学习优化器的思想,如归一化方法和动量方法。具体地说,该项目将开发一类新的预处理方法作为归一化方法的广泛适用的替代方法,并开发一类新的自适应动量方法作为固定动量方法的稳健替代方法。将建立相关的收敛理论,新方法将适用于最先进的神经网络结构,如变压器和图神经网络。在这个项目中开发的新算法旨在将数值分析中一些最有成效的想法带到神经网络优化的发展中。这个奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Deep learning is at the forefront of research in artificial intelligence and machine learning, impacting a variety of applications in data science such as computer vision, speech recognition, natural language processing, and bioinformatics. A key challenge in deep neural network learning is model optimization, which is used for network training. However, traditional optimization algorithms are not applicable, primarily due to the high complexity and nonlinearity of deep neural networks. The goal of this project is to develop novel robust optimization algorithms that can effectively address these difficulties and can more efficiently train deep learning models in practice. The project also involves the application of this work to the translation of equivalent chemical representations used in drug design as well as Bayesian inference for uncertainty quantification. As part of this project, graduate and undergraduate students will be trained in deep learning research, and software will be developed and made freely available.This project includes the development of two new classes of optimization algorithms that are built on the frameworks of traditional preconditioning and conjugate gradient methods but incorporate ideas from some successful specialized deep learning optimizers such as normalization methods and momentum methods. Specifically, the project will develop a new class of preconditioning methods as a widely applicable alternative to the normalization methods and a new class of adaptive momentum methods as a robust alternative to the fixed momentum methods. Related convergence theory will be established, and the new methods will be adapted to state-of-the-art neural network architectures such as transformer and graph neural networks. The novel algorithms developed in this project intend to bring some of the most fruitful ideas in numerical analysis to the advancement of neural network optimization.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1007/s00521-022-08168-3
发表时间: 2022-03
期刊: Neural Computing and Applications
影响因子: 6
作者: [K. D. G. Maduranga;Vasily Zadorozhnyy;Qiang Ye]
通讯作者: K. D. G. Maduranga;Vasily Zadorozhnyy;Qiang Ye
Improving Deep Neural Networks’ Training for Image Classification With Nonlinear Conjugate Gradient-Style Adaptive Momentum
使用非线性共轭梯度式自适应动量改进深度神经网络 - 图像分类训练
DOI: 10.1109/tnnls.2023.3255783
发表时间: 2023
期刊: IEEE Transactions on Neural Networks and Learning Systems
影响因子: 10.4
作者: [Wang, Bao, Ye, Qiang]
通讯作者: Ye, Qiang
DOI: 10.1016/j.jmapro.2023.03.011
发表时间: 2023-05
期刊: Journal of Manufacturing Processes
影响因子: 6.2
作者: [Rui Yu;Yue Cao;Heping Chen;Qiang Ye;Yuming Zhang]
通讯作者: Rui Yu;Yue Cao;Heping Chen;Qiang Ye;Yuming Zhang
DOI: 10.1021/acs.jcim.2c01526
发表时间: 2023-04
期刊: Journal of chemical information and modeling
影响因子: 5.6
作者: [Edison Mucllari;Vasily Zadorozhnyy;Qiang Ye;D. Nguyen]
通讯作者: Edison Mucllari;Vasily Zadorozhnyy;Qiang Ye;D. Nguyen
共 6 条
    RI: Small: Optimal Transport Generative Adversarial Networks: Theory, Algorithms, and Applications
    CDS&E: Efficient and Robust Recurrent Neural Networks
    Accurate Preconditioing for Computing Eigenvalues of Large and Extremely Ill-conditioned Matrices
    Collaborative Research: CDS&E-MSS: Robust Algorithms for Interpolation and Extrapolation in Manifold Learning
    海外基金