Optimizing Network Performance in Distributed Machine Learning

Optimizing Network Performance in Distributed Machine Learning
复制标题

DOI:
--
复制
发表时间:
2015-07
期刊:
--
影响因子:
--
通讯作者:
Luo Mai;C. Hong;Paolo Costa
Luo Mai;C. Hong;Paolo Costa
中科院分区:
其他
文献类型:
--
作者:
Luo Mai;C. Hong;Paolo Costa

文献摘要

被引文献

相似文献

为了应对不断增长的训练数据可用性,已经有几个建议将机器学习计算扩展到单个服务器之外,并将其分布在集群中。虽然这可以减少培训时间,但观察到的速度通常会受到网络瓶颈的限制。为了解决这个问题,我们设计了MLNET,一个基于主机的通信层,旨在提高分布式机器学习系统的网络性能。这是通过流量减少技术(以减少核心和边缘的网络负载)和流量管理(以减少平均培训时间)的组合来实现的。MLNET的一个关键特征是它与现有的硬件和软件基础设施兼容,因此可以立即部署。我们描述了支撑MLNET的主要技术,并通过仿真表明,总体训练时间可以减少78%。虽然我们的结果是初步的,但我们的结果表明了网络所起的关键作用以及引入新的通信层来提高分布式机器学习系统的性能的好处。
To cope with the ever growing availability of training data, there have been several proposals to scale machine learning computation beyond a single server and distribute it across a cluster. While this enables reducing the training time, the observed speed up is often limited by network bottlenecks. To address this, we design MLNET, a host-based communication layer that aims to improve the network performance of distributed machine learning systems. This is achieved through a combination of traffic reduction techniques (to diminish network load in the core and at the edges) and traffic management (to reduce average training time). A key feature of MLNET is its compatibility with existing hardware and software infrastructure so it can be immediately deployed. We describe the main techniques underpinning MLNET and show through simulation that the overall training time can be reduced by up to 78%. While preliminary, our results indicate the critical role played by the network and the benefits of introducing a new communication layer to increase the performance of distributed machine learning systems.