Collaborative Research: SHF: Medium: HERMES: On-Device Distributed Machine Learning via Model-Hardware Co-Design
Collaborative Research: SHF: Medium: HERMES: On-Device Distributed Machine Learning via Model-Hardware Co-Design
批准号:
2107085
负责人:
Diana Marculescu
金额:
$56.4万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-10-01 至 2024-09-30
中文摘要
机器学习(ML)有望成为现代社会最具颠覆性的技术,它将改变人类相互作用或与周围世界互动的方方面面。为了有效,ML模型必须使用大量数据,并且必须随时随地高效地构建和更新新的数据、设备或用户。为了满足消费者的需求或严格的设备或环境限制,ML系统必须快速响应并尽可能使用最少的能量,特别是在广泛分布的物联网(IoT)设备的背景下。该项目通过开发分布式培训的新方法来满足这一需求,该方法允许直接在物联网设备上进行现场快速且节能的培训。该项目的成果将直接影响广泛的应用,从人员流动性跟踪和预测,到实时语音或语言处理。此外,该项目旨在改变以多学科方式培训工程师的方式,以处理高效设计分布式ML系统的问题,这些系统对数据、设备或用户的可用性做出实时和低能源成本的响应。该项目旨在培养一批不同的研究实习生,同时扩大到高中和初中生群体。考虑到这项工作的统一跨学科方面、劳动力发展计划及其行业影响,该项目使新兴或成熟的工程师和行业合作伙伴能够进行广泛的协作。ML模型的大部分培训在云中集中完成,因此无法满足用户的隐私问题或响应时间,如果需要快速模型更新,则变得不适用。虽然高效的设备上推理一直是最近研究的热点,但设备上的分布式训练和推理尚未从响应时间和能效的角度进行处理;这对于物联网尤其重要,因为网络在训练和推理效率方面都发挥着重要作用。为了应对这些挑战,该项目(称为Hermes)提供了一种统一的多管齐下的方法,用于在设备上分布式设置中满足实时和能源限制。爱马仕确保ML方法和底层硬件是共同设计的,从而解决了当前的挑战,如私有数据共享、通信开销或分布式ML的实时和节能响应。更具体地说,Hermes包括:(I)一套可扩展的方法,用于硬件感知的实时、高能效的分布式训练,该方法基于联合学习和分布式优化,对数据和设备变异性具有健壮性;(Ii)ML模型和硬件的联合设计,包括利用硬件特征并识别满足约束的ML模型的超参数优化,以及高效地找到满足约束的架构的硬件设计探索;以及(Iii)用于展示所产生的ML系统的好处的分析和原型基础设施。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine Learning (ML) is poised to become the most disruptive technology in modern society by changing all aspects of how humans interact with each other or with the world around them. To be effective, ML models must use vast amounts of data and must be built and updated efficiently wherever and whenever new data, devices, or users are available. To satisfy consumer needs or stringent device or environmental constraints, ML systems must respond fast and use minimum energy whenever possible, especially in the context of widely spread Internet-of-Things (IoT) devices. This project addresses this need by developing new approaches for distributed training that allows for fast and energy efficient training in the field, directly on IoT devices. The results of this project are poised to directly impact a wide array of applications, ranging from human mobility tracking and prediction, to real-time speech or language processing. Furthermore, the project aims to change how engineers are trained in a multidisciplinary fashion for dealing with the problem of efficiently designing distributed ML systems that respond in real-time and with low energy cost to availability of data, devices, or users. The project aims to develop a body of diverse research trainees, while expanding outreach to high-school and middle-school student populations. Given the unified interdisciplinary aspects of this work, its workforce development plan, and its industrial impact, this project enables wide collaboration among emerging or established engineers and industrial partners.Most training of ML models is done centrally in the cloud, thereby not satisfying user privacy concerns or response times, and becoming inapplicable if fast model updates are needed. While efficient on-device inference has been an intense focus of recent research, on-device distributed training and inference have not been addressed from response time and energy efficiency perspectives; this is particularly important for IoT, where the network plays a major part both in training and inference efficiency. To address these challenges, this project (dubbed HERMES) provides a unified multipronged approach for meeting real-time and energy constraints in an on-device distributed setting. HERMES ensures that ML methods and underlying hardware are co-designed, thereby addressing current challenges of private data sharing, communication overhead, or real-time and energy-efficient response of distributed ML. More specifically, Hermes includes: (i) a set of scalable approaches for hardware-aware real-time, energy efficient distributed training based on federated learning and distributed optimization that is robust to data and device variability; (ii) the co-design of ML model and hardware, comprising hyperparameter optimization that exploits hardware characteristics and identifies constraint-satisfying ML models, and hardware design exploration that efficiently finds constraint satisfying architectures; and (iii) an analysis and prototyping infrastructure for demonstrating the benefits of resulting ML systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3576842.3582378
发表时间:
2023-05
期刊:
Proceedings of the 8th ACM/IEEE Conference on Internet of Things Design and Implementation
影响因子:
--
作者:
[Allen-Jasmin Farcas;Myungjin Lee;R. Kompella;Hugo Latapie;G. de Veciana;R. Marculescu]
通讯作者:
Allen-Jasmin Farcas;Myungjin Lee;R. Kompella;Hugo Latapie;G. de Veciana;R. Marculescu
Demo Abstract: A Hardware Prototype Targeting Federated Learning with User Mobility and Device Heterogeneity
演示摘要:针对具有用户移动性和设备异构性的联邦学习的硬件原型
DOI:
10.1145/3576842.3589160
发表时间:
2023
期刊:
IoTDI
影响因子:
--
作者:
[Farcas, Allen-Jasmin, Marculescu, Radu]
通讯作者:
Marculescu, Radu
MobileTL: On-Device Transfer Learning with Inverted Residual Blocks
MobileTL:具有倒置残差块的设备上迁移学习
DOI:
10.1609/aaai.v37i6.25874
发表时间:
2023
期刊:
Proceedings of the AAAI Conference on Artificial Intelligence
影响因子:
--
作者:
[Chiang, Hung-Yueh, Frumkin, Natalia, Liang, Feng, Marculescu, Diana]
通讯作者:
Marculescu, Diana
DOI:
10.1109/cvpr52729.2023.00682
发表时间:
2022-10
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
作者:
[Feng Liang;Bichen Wu;Xiaoliang Dai;Kunpeng Li;Yinan Zhao;Hang Zhang;Peizhao Zhang;Péter Vajda;D. Marculescu]
通讯作者:
Feng Liang;Bichen Wu;Xiaoliang Dai;Kunpeng Li;Yinan Zhao;Hang Zhang;Peizhao Zhang;Péter Vajda;D. Marculescu
DOI:
10.1109/tc.2023.3288755
发表时间:
2022-01
期刊:
IEEE Transactions on Computers
影响因子:
3.7
作者:
[Zihui Xue;Yuedong Yang-;Mengtian Yang;R. Marculescu]
通讯作者:
Zihui Xue;Yuedong Yang-;Mengtian Yang;R. Marculescu
共 8 条
Collaborative Research: CyberSEES: Climate-Aware Renewable Hydropower Generation and Disaster Avoidance
-
批准号:1331804
-
项目类别:Standard Grant
-
资助金额:$45.6万
-
财政年份:2013
-
负责人:Diana Marculescu
-
依托单位:
Planning Grant: I/UCRC for Nexys: Next Generation Electronic System Design
-
批准号:1160997
-
项目类别:Standard Grant
-
资助金额:$1.45万
-
财政年份:2012
-
负责人:Diana Marculescu
-
依托单位:
Collaborative Research: CSR---EHS: Cross-System Modeling and Management for Variation-Adaptive Computing
-
批准号:0720529
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2007
-
负责人:Diana Marculescu
-
依托单位:
CSR---SMA: Variability-Aware System Level Performance and Power Analysis
-
批准号:0720653
-
项目类别:Continuing Grant
-
资助金额:$29.0万
-
财政年份:2007
-
负责人:Diana Marculescu
-
依托单位:
Variability-Energy Interactions at the Microarchitecture to System-Level Interface for 2D and 3D Architectures
-
批准号:0702451
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Diana Marculescu
-
依托单位:
SGER: Analysis of Fault-Tolerant Nanoscale Designs
-
批准号:0542644
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Diana Marculescu
-
依托单位:
CAREER: Software Level Power Analysis and Optimization
-
批准号:0084479
-
项目类别:Continuing Grant
-
资助金额:$26.0万
-
财政年份:2000
-
负责人:Diana Marculescu
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: