Performance, Portability, and Productivity for Deep Learning Applications on Multi- and Many-Core Architectures (PPP-DL)
Performance, Portability, and Productivity for Deep Learning Applications on Multi- and Many-Core Architectures (PPP-DL)
批准号:
470527619
负责人:
Professor Dr. Sergei Gorlatch
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:
中文摘要
深度学习是目前最流行的机器学习方法,可以解决学术界和工业界的各种实际问题。数字图书馆应用程序的成功关键取决于为多核CPU、图形处理单元(GPU)、现场可编程门阵列(FPGA)等现代并行体系结构实现数字图书馆算法的软件的质量。TensorFlow和PyTorch等最先进的数字图书馆框架依赖于供应商提供的通用库(如Intel或NVIDIA)的高性能,导致三个基本方面的主要弱点:i)次优性能-由于库侧重于通用用途,许多特定于数字图书馆的优化不适用;Ii)缺乏功能和性能可移植性,因为库仅针对特定供应商的架构进行了专门设计和优化;iii)限制了用户生产力,因为库仅限于一组固定的预先实现的算法(例如,矩阵乘法和卷积),并且将高性能库集成到DL框架中是繁琐的。该项目将为面向现代并行架构的数字图书馆应用程序开发一种新颖的、全面的自动代码生成和优化方法;其总体目标是在一个组合方法中解决数字图书馆高性能计算领域的三个主要研究挑战:性能、可移植性和生产力(PPP)。我们计划在以下新贡献的基础上实现项目的目标:1)新的代数形式和基于形式的领域特定语言(DSL),用于在高层抽象方便地表达/实现已建立的和新兴的领域特定语言(DSL),从而有助于程序员的生产力;2)用于DL应用的统一的低级编程模型,该模型通过在实践状态的并行编程方法:OpenMP、CUDA、OpenCL等中直接将代码降低为可执行代码来实现代码的功能可移植性;3)我们DSL的代码生成机制,它通过在我们的低级编程模型中自动生成可自动调整的代码,在各种体系结构和输入/输出特性上实现高、可移植的性能;4)基于新兴的MLIR框架,将我们的代码生成机制集成到现代DL框架中的系统过程;5)新的自动调整系统,它通过组合的数字搜索技术完全自动优化我们生成的代码;6)一种新的分析成本模型,用于预测不同体系结构的动态链接库应用程序的运行时间,以加快自动调优过程。我们将在各种动态链接库应用程序、并行体系结构和真实世界的动态链接库数据集的所有性能、可移植性和生产率方面对我们的方法进行实验比较。
英文摘要
Deep Learning (DL) is currently the most popular machine-learning method that solves a great variety of real-world problems in academia and industry. The success of DL applications critically depends on the quality of software that implements DL algorithms for modern parallel architectures like multi-core CPU, Graphics Processing Unit (GPU), Field-Programmable Gate Array (FPGA), etc. The state-of-the-art DL frameworks like TensorFlow and PyTorch rely for high performance upon general-purpose libraries provided by vendors, such as Intel or NVIDIA, causing major weaknesses regarding three fundamental aspects: i) suboptimal performance – many DL-specific optimizations are not applicable because of libraries’ focus toward general-purpose usage; ii) lacking both functional and performance portability, because the libraries are specifically designed and optimized toward architectures of particular vendors only; iii) restricted user productivity, because the libraries are limited to a fixed set of pre- implemented algorithms (e.g., matrix multiplication and convolutions), and it is cumbersome to integrate high-performance libraries into DL frameworks. This project will develop a novel, holistic approach toward automatic code generation and optimization for DL applications targeting modern parallel architectures; its overall goal is to address in one combined approach three major research challenges in the area of high-performance computing for DL: Performance, Portability, and Productivity (PPP). We plan to achieve the goal of the project based on the following new contributions: 1) a new algebraic formalism and a formalism-based Domain-Specific Language (DSL) for conveniently expressing/implementing established and emerging DL applications at a high-level of abstraction, thereby contributing to programmer’s productivity; 2) a uniform low-level programming model for DL applications, which enables functional portability of code by being straightforwardly lowerable to executable code in the state-of-practice parallel programming approaches: OpenMP, CUDA, OpenCL, etc.; 3) a code generation mechanism for our DSL that enables high, portable performance over various architectures and input/output characteristics by automatically generating auto-tunable code in our low-level programming model; 4) a systematic process that integrates our code generation mechanism into modern DL frameworks, based on the emerging MLIR framework; 5) a new auto-tuning system that fully automatically optimizes our generated code via combined numerical search techniques; 6) a new analytical cost model to predict for different architectures the run time of DL applications expressed in our DSL, in order to accelerate the auto-tuning process.We will experimentally compare our approach in terms of all – performance, portability, and productivity – to state-of-the-art approaches for a broad range of DL applications, parallel architectures, and real-world DL data sets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collective Operations: Formal Framework, Equalities, Efficiency
-
批准号:5264766
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2000
-
负责人:Professor Dr. Sergei Gorlatch
-
依托单位:
海外基金