Matching DNN Compression and Cooperative Training with Resources and Data Availability

Matching DNN Compression and Cooperative Training with Resources and Data Availability
复制标题

DOI:
10.1109/infocom53939.2023.10229076
复制
发表时间:
2022-12
期刊:
IEEE INFOCOM 2023 - IEEE Conference on Computer Communications
影响因子:
--
通讯作者:
F. Malandrino;G. Giacomo;Armin Karamzade;M. Levorato;C. Chiasserini
F. Malandrino;G. Giacomo;Armin Karamzade;M. Levorato;C. Chiasserini
中科院分区:
其他
文献类型:
--
作者:
F. Malandrino;G. Giacomo;Armin Karamzade;M. Levorato;C. Chiasserini

文献摘要

被引文献

相似文献

为了使机器学习(ML)可持续并易于在相关数据所在的各种设备上运行,必须根据需要压缩ML模型,同时仍然满足所需的学习质量和时间性能。然而,ML模型应该压缩多少以及何时压缩,以及应该在哪里执行其训练,这是很难做出的决定,因为它们取决于模型本身,可用节点的资源以及这些节点拥有的数据。现有的研究集中在这些方面的每一个单独的,但是,他们没有考虑到这些决定如何可以共同作出和相互适应。在这项工作中,我们对专注于DNN训练的网络系统进行建模,将上述多维问题形式化,并考虑到其NP难度,制定了一个近似的动态规划问题,我们通过PACT算法框架解决该问题。重要的是,PACT利用了一个代表学习过程的时间扩展图,以及一种数据驱动和理论方法来预测训练决策所导致的损失演变。我们证明,PACT的解决方案可以得到尽可能接近所需的最佳,在成本增加的时间复杂性,而且,在任何情况下,这种复杂性是多项式。数值结果还表明,即使在最不利的设置,PACT优于国家的最先进的替代品,密切配合的最佳能源成本。
To make machine learning (ML) sustainable and apt to run on the diverse devices where relevant data is, it is essential to compress ML models as needed, while still meeting the required learning quality and time performance. However, how much and when an ML model should be compressed, and where its training should be executed, are hard decisions to make, as they depend on the model itself, the resources of the available nodes, and the data such nodes own. Existing studies focus on each of those aspects individually, however, they do not account for how such decisions can be made jointly and adapted to one another. In this work, we model the network system focusing on the training of DNNs, formalize the above multi-dimensional problem, and, given its NP-hardness, formulate an approximate dynamic programming problem that we solve through the PACT algorithmic framework. Importantly, PACT leverages a time-expanded graph representing the learning process, and a data-driven and theoretical approach for the prediction of the loss evolution to be expected as a consequence of training decisions. We prove that PACT’s solutions can get as close to the optimum as desired, at the cost of an increased time complexity, and that, in any case, such complexity is polynomial. Numerical results also show that, even under the most disadvantageous settings, PACT outperforms state-of-the-art alternatives and closely matches the optimal energy cost.