课题基金 / 基金详情

Training Unstructured Sparse Neural Networks

Training Unstructured Sparse Neural Networks
训练非结构化稀疏神经网络
批准号:
RGPIN-2022-03120
负责人:
Ioannou, Yani
金额:
$1.82万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Ioannou, Yani的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The Increasing Cost of Training Deep Neural Networks Deep Neural Networks (DNNs) are behind the intelligence found in contemporary technology, enabling us to search our photos, ask questions of smart assistants or use machine translation to understand a foreign language. DNNs have a fundamental problem however: they are very expensive - in cost, energy usage, and time, both during training (learning a task) and inference (application). The state-of-the-art model for Natural Language Processing (NLP), GPT-3, has 175 billion parameters and is estimated to cost more than $4.6M USD to train. The current trend is that training DNNs will only get ever more expensive as the growth in the cost of training state-of-the-art DNN models has surpassed even the exponential growth in transistor technology that governs our computational capacity, i.e. Moore's Law. Proposed Program of Research Enable Sparse DNN Training DNNs are primarily expensive due to the empirically established, but poorly understood, requirement to over-parameterize DNNs models during training for good generalization (performance on unseen data). We know these models are over-parameterized because, after training, 80-95% of the learned weights (parameters) of DNN models can be removed (pruned) without significant loss of generalization [16]. Unstructured pruning, that is removing unnecessary individual weights from a DNN, can be highly effective at reducing the size and efficiency of pre-trained DNNs used in applications. However, attempting to train unstructured sparse DNN from random initialization - just as dense (standard) DNNs are trained - rarely achieves similar generalization as dense training empirically. For this reason, unstructured sparsity has not played a significant role in decreasing the cost of training DNNs as of yet. Developing effective approaches for efficient training would make DNN training cheaper, faster, more repeatable, and importantly more accessible to both researchers and new applications. Efficient Deep Learning and Society: Adversarial Robustness and Bias of Efficient DNNs Already DNNs used in industrial application are drastically different from those proposed in most academic papers, in using efficient DNNs methods (e.g. quantization, pruning, and distillation) and efficient DNN architectures - and yet, there is little work exploring the differences between the well-studied academic models, and industry-applied efficient DNNs beyond simply generalization performance and efficiency. Understanding the differences between efficient DNNs and the typical DNNs seen in academic research is becoming imperative given the increasing reliance of our society on this technology. With an ever larger potential to affect our society, identifying and resolving any issues with robustness and bias specific to efficient DNNs is an increasingly important area of research, and one relevant to a variety of potential industrial partners using DNNs in real-world applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Training Unstructured Sparse Neural Networks
  • 批准号:
    DGECR-2022-00358
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2022
  • 负责人:
    Ioannou, Yani
  • 依托单位:
Training Unstructured Sparse Neural Networks
  • 批准号:
    DGDND-2022-03120
  • 项目类别:
    DND/NSERC Discovery Grant Supplement
  • 资助金额:
    $2.91万
  • 财政年份:
    2022
  • 负责人:
    Ioannou, Yani
  • 依托单位:
海外基金