Dynamic Network-Assisted D2D-Aided Coded Distributed Learning

Dynamic Network-Assisted D2D-Aided Coded Distributed Learning
复制标题

DOI:
10.1109/tcomm.2023.3259442
复制
发表时间:
2021-11
影响因子:
8.3
通讯作者:
Nikita Zeulin;O. Galinina;N. Himayat;Sergey D. Andreev;R. Heath
Nikita Zeulin;O. Galinina;N. Himayat;Sergey D. Andreev;R. Heath
中科院分区:
计算机科学2区
文献类型:
--
作者:
Nikita Zeulin;O. Galinina;N. Himayat;Sergey D. Andreev;R. Heath

文献摘要

相似文献

如今,许多机器学习(ML)应用程序在无线网络边缘提供连续数据处理和实时数据分析。分布式实时ML解决方案非常容易受到由资源异构性引起的所谓的离散效应的影响,这可以通过各种计算卸载机制来缓解,这些机制严重影响通信效率,特别是在大规模场景中。为了减少通信开销,我们利用设备到设备(D2 D)连接,这提高了频谱利用率,并允许邻近设备之间的有效数据交换。特别是,我们设计了一种新的D2 D辅助编码的分布式学习方法命名为D2 D-CDL跨设备的高效负载平衡。所提出的解决方案捕获系统动态,包括数据(时变学习模型,数据到达的不规则强度),设备(不同的计算资源和训练数据量)和部署(不同的位置和D2 D图形连接)。为了减少通信轮数,我们推导出一个最佳的压缩率,最大限度地减少处理时间。由此产生的优化问题提供了次优的压缩参数,提高了总的训练时间。我们提出的方法对于实时协作应用程序特别有益,在这些应用程序中,用户不断生成训练数据。
Today, numerous machine learning (ML) applications offer continuous data processing and real-time data analytics at the edge of wireless networks. Distributed real-time ML solutions are highly susceptible to the so-called straggler effect caused by resource heterogeneity, which can be mitigated by various computation offloading mechanisms that severely impact communication efficiency, especially in large-scale scenarios. To reduce the communication overhead, we leverage device-to-device (D2D) connectivity, which enhances spectrum utilization and allows for efficient data exchange between proximate devices. In particular, we design a novel D2D-aided coded distributed learning method named D2D-CDL for efficient load balancing across devices. The proposed solution captures system dynamics, including data (time-varying learning model, irregular intensity of data arrivals), device (diverse computational resources and volume of training data), and deployment (different locations and D2D graph connectivity). To decrease the number of communication rounds, we derive an optimal compression rate, which minimizes the processing time. The resulting optimization problem provides suboptimal compression parameters that improve the total training time. Our proposed method is particularly beneficial for real-time collaborative applications, where users continuously generate training data.