课题基金 / 基金详情

Towards Improved Understanding and Efficient Utilization of Depthwise Computation in Modern Neural Networks

Towards Improved Understanding and Efficient Utilization of Depthwise Computation in Modern Neural Networks
提高对现代神经网络深度计算的理解和有效利用
批准号:
577088-2022
负责人:
Wolf, GuyG
金额:
$3.28万
依托单位:
依托单位国家:
加拿大
项目类别:
Alliance Grants
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Deep Neural Networks have shown remarkable performance in a wide array of tasks including image recognition, speech recognition, and text generation. They consist of a number of sequential modules typically jointly optimized for a specific task. The functional behavior of the individual modules that emerges in such systems is complex and difficult to understand or analyze, while at the same time yielding remarkable results. On the other hand, the computational and memory efficiency of deep networks is a large bottleneck to their widespread usability. In this project we propose to tackle these problems in a holistic manner. In the first step, we propose to develop and enhance analytical tools for comparing and analyzing intermediate representations of deep networks. Borrowing methods from kernel non-parametrics, our aim is to provide analytical tools to determine how representations evolve in depth, and what information is extracted at each layer. Ultimately we would like to use these tools to characterize an implicit objective function that emerges in depth. These aspects will be studied in both traditional feed-forward networks and modern networks with skip connections, as well as Transformer-based models utilizing self-attention layers. In the second step we will use these techniques to enhance local learning methods as part of principled local objective functions, targeted to allow increasing parallelism and saving memory when training large deep neural networks. Concurrently, we will investigate continuous-depth deep neural networks. Our work will utilize a variety of modern deep neural networks with a focus on the emerging Transformer architecture. Finally, we will apply the developed techniques in a broad set of application areas including natural language processing, computer vision, drug discovery, and recommender systems. Development of this will require interdisciplinary collective expertise in theoretical tools, model-parallel learning, and continuous depth networks. Over this three-year project the PI and co-investigators will train three PhD students and one MSc student across a wide range of industrially relevant skills, including deep learning and distributed learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金