Performance, Portability, and Productivity for Deep Learning Applications on Multi- and Many-Core Architectures (PPP-DL)
Performance, Portability, and Productivity for Deep Learning Applications on Multi- and Many-Core Architectures (PPP-DL)
批准号:
470527619
负责人:
Professor Dr. Sergei Gorlatch
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:
中文摘要
深度学习(DL)是目前最流行的机器学习方法,它解决了学术界和工业界各种各样的现实问题。深度学习应用的成功在很大程度上取决于为现代并行架构(如多核CPU、图形处理单元(GPU)、现场可编程门阵列(FPGA)等)实现深度学习算法的软件质量。最先进的深度学习框架,如TensorFlow和PyTorch,依赖于供应商(如Intel或NVIDIA)提供的通用库来实现高性能,这在三个基本方面造成了主要的弱点:i)次优性能——许多特定于DL的优化不适用,因为库的重点是通用用途;Ii)缺乏功能和性能的可移植性,因为这些库是专门针对特定供应商的架构设计和优化的;iii)限制用户生产力,因为库仅限于一组固定的预实现算法(例如,矩阵乘法和卷积),并且将高性能库集成到DL框架中很麻烦。该项目将开发一种新颖的整体方法,用于针对现代并行架构的深度学习应用程序的自动代码生成和优化;它的总体目标是用一种结合的方法解决DL高性能计算领域的三个主要研究挑战:性能、可移植性和生产力(PPP)。我们计划基于以下新贡献来实现项目的目标:1)一个新的代数形式化和基于形式化的领域特定语言(DSL),用于方便地在抽象的高层表达/实现已建立和新兴的DL应用程序,从而有助于程序员的生产力;2)为DL应用程序提供统一的低级编程模型,通过直接降低代码的可执行性,实现代码的功能可移植性,在实践状态的并行编程方法中:OpenMP, CUDA, OpenCL等;3) DSL的代码生成机制,通过在低级编程模型中自动生成自动可调代码,实现各种架构和输入/输出特性的高可移植性能;4)基于新兴的MLIR框架,将我们的代码生成机制集成到现代DL框架中的系统流程;5)一个新的自动调谐系统,通过组合数值搜索技术全自动优化我们生成的代码;6)建立了一种新的分析成本模型,用于预测不同架构的DSL中表示的DL应用程序的运行时间,以加快自动调优过程。我们将通过实验比较我们的方法在所有性能、可移植性和生产力方面与最先进的方法,用于广泛的深度学习应用程序、并行架构和现实世界的深度学习数据集。
英文摘要
Deep Learning (DL) is currently the most popular machine-learning method that solves a great variety of real-world problems in academia and industry. The success of DL applications critically depends on the quality of software that implements DL algorithms for modern parallel architectures like multi-core CPU, Graphics Processing Unit (GPU), Field-Programmable Gate Array (FPGA), etc. The state-of-the-art DL frameworks like TensorFlow and PyTorch rely for high performance upon general-purpose libraries provided by vendors, such as Intel or NVIDIA, causing major weaknesses regarding three fundamental aspects: i) suboptimal performance – many DL-specific optimizations are not applicable because of libraries’ focus toward general-purpose usage; ii) lacking both functional and performance portability, because the libraries are specifically designed and optimized toward architectures of particular vendors only; iii) restricted user productivity, because the libraries are limited to a fixed set of pre- implemented algorithms (e.g., matrix multiplication and convolutions), and it is cumbersome to integrate high-performance libraries into DL frameworks. This project will develop a novel, holistic approach toward automatic code generation and optimization for DL applications targeting modern parallel architectures; its overall goal is to address in one combined approach three major research challenges in the area of high-performance computing for DL: Performance, Portability, and Productivity (PPP). We plan to achieve the goal of the project based on the following new contributions: 1) a new algebraic formalism and a formalism-based Domain-Specific Language (DSL) for conveniently expressing/implementing established and emerging DL applications at a high-level of abstraction, thereby contributing to programmer’s productivity; 2) a uniform low-level programming model for DL applications, which enables functional portability of code by being straightforwardly lowerable to executable code in the state-of-practice parallel programming approaches: OpenMP, CUDA, OpenCL, etc.; 3) a code generation mechanism for our DSL that enables high, portable performance over various architectures and input/output characteristics by automatically generating auto-tunable code in our low-level programming model; 4) a systematic process that integrates our code generation mechanism into modern DL frameworks, based on the emerging MLIR framework; 5) a new auto-tuning system that fully automatically optimizes our generated code via combined numerical search techniques; 6) a new analytical cost model to predict for different architectures the run time of DL applications expressed in our DSL, in order to accelerate the auto-tuning process.We will experimentally compare our approach in terms of all – performance, portability, and productivity – to state-of-the-art approaches for a broad range of DL applications, parallel architectures, and real-world DL data sets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collective Operations: Formal Framework, Equalities, Efficiency
-
批准号:5264766
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2000
-
负责人:Professor Dr. Sergei Gorlatch
-
依托单位:
海外基金