Monitoring and Debugging of High Performance Distributed Heterogeneous Cloud Applications
Monitoring and Debugging of High Performance Distributed Heterogeneous Cloud Applications
批准号:
554158-2020
负责人:
Dagenais, Michel
金额:
$24.98万
依托单位国家:
加拿大
项目类别:
Alliance Grants
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
通信和计算基础设施正以极快的速度变得越来越复杂。最近的应用包括5G连接的移动设备、自动驾驶汽车、智能机器人和由机器学习驱动的智能数字助理。这些进步是由硬件和软件层面的许多技术发展实现的,例如具有数十个核心的计算机中央处理单元,具有数千个核心和超过180亿个逻辑元素的图形和密集计算(GPGPU)协处理器,5G低延迟高速网络以及并行执行请求的基于云的基础设施。因此,即使是发起电话呼叫、进行Web搜索、路由数据包或显示视频帧等简单操作,也可能涉及多个处理单元(可能是多个服务器)上的多个并行核。此外,同样的操作,在几秒钟后,可能会由云中的不同核心和物理服务器以不同的方式提供服务。因此,了解这些操作的性能变得极其困难,并且严重缺乏用于此目的的工具。在本项目中,将扩展高性能分布式系统的跟踪、分析、调试和监控工具,以有效地从所有层的所有单元中提取信息,从硬件到应用程序,并应对大量的内核和计算机。该项目特别关注通过边缘服务器和5G网络连接到移动和物联网设备的云应用程序,利用新一代共享内存gpgpu的高性能计算,机器学习应用程序以及用于更集成的软件开发工具的新模块化架构。因此,高性能分布式系统的设计人员和操作人员将拥有快速分析系统性能的工具,自动或手动发现问题,并优化操作。# (cr) #(低频)L 'infrastructure计算et de沟通上非常很快et devient始终+ sophistiquee。新应用的例子包括l'informatique移动通信设备、s ' communications communications 5G、s ' voitures autonomes、s ' robots et assistant、s ‘ communications智能和s ’学徒机。从不同的角度看,有不同的可能性;从不同的角度看,有不同的可能性;从不同的角度看,有不同的可能性;从不同的角度看,有不同的可能性;从不同的角度看,有不同的可能性;从不同的角度看,有不同的可能性;从不同的角度看,有不同的可能性。一个简单的requête, comme faire one recherche sur la Toile,或者comme faire one image d'un vidsamoise,或者comme faire one imageche sur vidsamoise,或者comme faire one imageche sur vidsamoise。加上encore, la même openciration, quelques secondes Plus ard, pourrait être servie de maniires diffenente, par diffencients coeures et noeuds physiques, dans l'infonuagique。在印度,理解那些不合格的人的表现extrêmement困难,以及那些不合格的人的表现。项目管理、项目管理、概况、项目管理、项目管理、项目管理、项目管理、项目管理、项目管理、项目管理、项目管理、项目管理、项目管理、项目管理、项目管理、项目管理。将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起,将项目集中在一起。从概念上讲,从系统上讲,从系统上讲,从系统上讲,从系统上讲,从系统上讲,从系统上讲,从系统上讲,从系统上讲,从系统上讲
英文摘要
The communication and computing infrastructure is getting ever more sophisticated at an extremely rapid pace. Recent applications include 5G connected mobile devices, autonomous cars, smart robots and intelligent digital assistants powered by Machine Learning. These advances are made possible by a number of technological developments at the hardware and software levels, such as computer central processing units with tens of cores, coprocessors for graphics and intensive computations (GPGPU) with thousands of cores and over 18 billion logic elements, 5G low latency high speed networking, and Cloud based infrastructures that execute their requests in parallel.As a result, even a simple operation such as initiating a phone call, making a Web search, routing a packet or displaying a video frame, can involve many parallel cores on more than one processing unit, possibly on several servers. Moreover, the same operation, a few seconds later, may be served in a different way by different cores and physical servers in the Cloud. Therefore, understanding the performance of these operations has become extremely difficult and the tools for that purpose are severely lacking. In this project, the tracing, profiling, debugging and monitoring tools for High Performance Distributed Systems will be extended to efficiently extract information from all units in all layers, from the hardware to the applications, and cope with the large number of cores and computers. The project has a specific focus on Cloud applications connecting to mobile and Internet of Things devices through Edge servers and 5G networks, High Performance Computing exploiting the new generation of shared memory GPGPUs, Machine Learning applications, and a new modular architecture for more integrated software development tools. As a result, the designers and operators of High Performance Distributed Systems will have the tools in hand to quickly analyse their system performance, automatically or manually find problems, and optimise operations.#(cr)#(lf)L'infrastructure de calcul et de communication évolue très rapidement et devient toujours plus sophistiquée. Des exemples de nouvelles applications incluent l'informatique mobile avec les réseaux 5G, les voitures autonomes, les robots et assistants numériques intelligents et l'apprentissage machine. Ceci devient possible grâce à plusieurs avancées technologiques, tant matérielles que logicielles, comme les unités centrales de traitement avec des dizaines de coeurs, les coprocesseurs pour le graphisme et les calculs scientifiques (GPGPU) avec des milliers de coeurs, les réseaux 5G à haute vitesse et faible latence, et les infrastructures infonuagiques qui exécutent un grand nombre de requêtes en parallèle.En conséquence, une simple requête, comme faire une recherche sur la Toile ou afficher une image d'un vidéo, peut impliquer un grand nombre de coeurs en parallèle, possiblement sur plusieurs serveurs. Plus encore, la même opération, quelques secondes plus tard, pourrait être servie de manière différente, par différents coeurs et noeuds physiques, dans l'infonuagique. Ainsi, comprendre la performance de telles opérations devient extrêmement difficile et les outils disponibles présentent de nombreuses lacunes. Dans le cadre de ce projet, les outils pour le traçage, profilage, débogage et monitoring de systèmes répartis seront étendus pour efficacement extraire l'information de toutes les unités, à tous les niveaux, du matériel jusqu'aux applications. Le projet se concentre plus spécifiquement sur les applications infonuagiques connectant l'Internet des Objets à travers les réseaux 5G, l'utilisation des processeurs hétérogènes de type GPGPU pour les applications de haute performance, les applications d'apprentissage machine, et les architectures modulaires pour les environnements intégrés de développement. Ceci permettra aux concepteurs et opérateurs de systèmes répartis de haute performance d'avoir les outils en main pour rapidement
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Reinventing the tuning and debugging tools for multi-thousand cores computer systems
-
批准号:RGPIN-2017-05634
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2022
-
负责人:Dagenais, Michel
-
依托单位:
Reinventing the tuning and debugging tools for multi-thousand cores computer systems
-
批准号:RGPIN-2017-05634
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2021
-
负责人:Dagenais, Michel
-
依托单位:
Automated monitoring and debugging of large scale manycore heterogeneous systems
-
批准号:507883-2016
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$9.25万
-
财政年份:2020
-
负责人:Dagenais, Michel
-
依托单位:
Monitoring and Debugging of High Performance Distributed Heterogeneous Cloud Applications
-
批准号:554158-2020
-
项目类别:Alliance Grants
-
资助金额:$24.98万
-
财政年份:2020
-
负责人:Dagenais, Michel
-
依托单位:
Reinventing the tuning and debugging tools for multi-thousand cores computer systems
-
批准号:RGPIN-2017-05634
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2020
-
负责人:Dagenais, Michel
-
依托单位:
Automated monitoring and debugging of large scale manycore heterogeneous systems
-
批准号:507883-2016
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$18.5万
-
财政年份:2019
-
负责人:Dagenais, Michel
-
依托单位:
Reinventing the tuning and debugging tools for multi-thousand cores computer systems
-
批准号:RGPIN-2017-05634
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2019
-
负责人:Dagenais, Michel
-
依托单位:
Reinventing the tuning and debugging tools for multi-thousand cores computer systems
-
批准号:RGPIN-2017-05634
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2018
-
负责人:Dagenais, Michel
-
依托单位:
Automated monitoring and debugging of large scale manycore heterogeneous systems
-
批准号:507883-2016
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$18.5万
-
财政年份:2018
-
负责人:Dagenais, Michel
-
依托单位:
Automated monitoring and debugging of large scale manycore heterogeneous systems
-
批准号:507883-2016
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$9.25万
-
财政年份:2017
-
负责人:Dagenais, Michel
-
依托单位:
Reinventing the tuning and debugging tools for multi-thousand cores computer systems
-
批准号:RGPIN-2017-05634
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2017
-
负责人:Dagenais, Michel
-
依托单位:
Probing and Monitoring live distributed many-core systems
-
批准号:36677-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2016
-
负责人:Dagenais, Michel
-
依托单位:
Software debugging and monitoring for heterogeneous many-core telecom systems
-
批准号:468687-2014
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$10.2万
-
财政年份:2015
-
负责人:Dagenais, Michel
-
依托单位:
Probing and Monitoring live distributed many-core systems
-
批准号:36677-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2015
-
负责人:Dagenais, Michel
-
依托单位:
Integrated tracing, profiling and debugging for tuning large heterogeneous clusters
-
批准号:424666-2011
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$10.64万
-
财政年份:2014
-
负责人:Dagenais, Michel
-
依托单位:
Probing and Monitoring live distributed many-core systems
-
批准号:36677-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2014
-
负责人:Dagenais, Michel
-
依托单位:
Software debugging and monitoring for heterogeneous many-core telecom systems
-
批准号:468687-2014
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$10.2万
-
财政年份:2014
-
负责人:Dagenais, Michel
-
依托单位:
Probing and Monitoring live distributed many-core systems
-
批准号:36677-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2013
-
负责人:Dagenais, Michel
-
依托单位:
Diagnostics for real time distributed multi-core architecture in avionics
-
批准号:419922-2011
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$3.64万
-
财政年份:2013
-
负责人:Dagenais, Michel
-
依托单位:
Online surveillance of critical computer systems through advanced host-based detection
-
批准号:416802-2011
-
项目类别:Department of National Defence / NSERC Research Partnership
-
资助金额:$11.12万
-
财政年份:2013
-
负责人:Dagenais, Michel
-
依托单位:
海外基金