Towards Improved Understanding and Efficient Utilization of Depthwise Computation in Modern Neural Networks
Towards Improved Understanding and Efficient Utilization of Depthwise Computation in Modern Neural Networks
批准号:
577088-2022
负责人:
Wolf, GuyG
金额:
$3.28万
依托单位:
依托单位国家:
加拿大
项目类别:
Alliance Grants
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
深度神经网络在包括图像识别、语音识别和文本生成在内的广泛任务中表现出卓越的性能。它们由许多顺序模块组成,通常是针对特定任务联合优化的。在这样的系统中出现的单个模块的功能行为是复杂的,难以理解或分析,同时产生显著的结果。另一方面,深度网络的计算和存储效率是其广泛可用性的一大瓶颈。在这个项目中,我们建议以整体的方式解决这些问题。在第一步,我们建议开发和增强分析工具,用于比较和分析深度网络的中间表示。借用核非参数的方法,我们的目标是提供分析工具来确定表征如何深入发展,以及在每一层提取什么信息。最终,我们希望使用这些工具来描述一个隐含的目标函数。这些方面将在传统的前馈网络和具有跳过连接的现代网络以及利用自关注层的基于变压器的模型中进行研究。在第二步中,我们将使用这些技术来增强局部学习方法,作为原则局部目标函数的一部分,目的是在训练大型深度神经网络时增加并行性并节省内存。同时,我们将研究连续深度深度神经网络。我们的工作将利用各种现代深度神经网络,重点关注新兴的Transformer架构。最后,我们将开发的技术应用于广泛的应用领域,包括自然语言处理,计算机视觉,药物发现和推荐系统。这方面的发展需要跨学科的集体专业知识,包括理论工具、模型并行学习和持续深度网络。在这个为期三年的项目中,PI和合作研究人员将培训三名博士研究生和一名硕士研究生,涉及广泛的工业相关技能,包括深度学习和分布式学习。
英文摘要
Deep Neural Networks have shown remarkable performance in a wide array of tasks including image recognition, speech recognition, and text generation. They consist of a number of sequential modules typically jointly optimized for a specific task. The functional behavior of the individual modules that emerges in such systems is complex and difficult to understand or analyze, while at the same time yielding remarkable results. On the other hand, the computational and memory efficiency of deep networks is a large bottleneck to their widespread usability. In this project we propose to tackle these problems in a holistic manner. In the first step, we propose to develop and enhance analytical tools for comparing and analyzing intermediate representations of deep networks. Borrowing methods from kernel non-parametrics, our aim is to provide analytical tools to determine how representations evolve in depth, and what information is extracted at each layer. Ultimately we would like to use these tools to characterize an implicit objective function that emerges in depth. These aspects will be studied in both traditional feed-forward networks and modern networks with skip connections, as well as Transformer-based models utilizing self-attention layers. In the second step we will use these techniques to enhance local learning methods as part of principled local objective functions, targeted to allow increasing parallelism and saving memory when training large deep neural networks. Concurrently, we will investigate continuous-depth deep neural networks. Our work will utilize a variety of modern deep neural networks with a focus on the emerging Transformer architecture. Finally, we will apply the developed techniques in a broad set of application areas including natural language processing, computer vision, drug discovery, and recommender systems. Development of this will require interdisciplinary collective expertise in theoretical tools, model-parallel learning, and continuous depth networks. Over this three-year project the PI and co-investigators will train three PhD students and one MSc student across a wide range of industrially relevant skills, including deep learning and distributed learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金