Distributed Learning of Fully Connected Neural Networks using Independent Subnet Training

Distributed Learning of Fully Connected Neural Networks using Independent Subnet Training
复制标题

DOI:
10.14778/3529337.3529343
复制
发表时间:
2019-10
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Binhang Yuan;Anastasios Kyrillidis;C. Jermaine
Binhang Yuan;Anastasios Kyrillidis;C. Jermaine
中科院分区:
其他
文献类型:
--
作者:
Binhang Yuan;Anastasios Kyrillidis;C. Jermaine

文献摘要

相似文献

分布式机器学习(ML)可以带来比单机学习更多的计算资源来承受的,从而可以减少培训时间。分布式学习分区模型和许多机器上的数据,允许模型和数据集大小超出单个计算机的可用计算功率和内存。但是,实际上,分布式ML在强制性分配时具有挑战性,而不是由从业者选择。在这种情况下,由于每个工人的内存能力有限,甚至由于数据隐私问题,因此不可避免地会在工人中分离数据。在那里,由于工人之间的主导转移成本,现有的分布式方法将完全失败,甚至不适用。我们提出了一种新的方法来分布完全连接的神经网络学习,称为独立子网培训(IST),以处理这些情况。在IST中,原始网络分解为具有相同深度的一组狭窄子网。然后,在交换参数以产生新的子网和训练周期重复之前,对这些子网进行本地训练。这种自然的“模型并行”方法仅通过在每个设备上存储一部分网络参数来限制内存使用量。此外,没有任何要求工人之间共享数据(即子网培训是本地和独立的),并且通过将原始网络分解为独立子网,可以降低通信量和频率。 IST的这些属性可以应对由于分布式数据,较慢的互连或有限的设备内存而应对的问题,这使IST成为强制性分布案例的合适方法。我们通过实验表明,IST导致培训时间远低于共同的分布式学习方法。
Distributed machine learning (ML) can bring more computational resources to bear than single-machine learning, thus enabling reductions in training time. Distributed learning partitions models and data over many machines, allowing model and dataset sizes beyond the available compute power and memory of a single machine. In practice though, distributed ML is challenging when distribution is mandatory, rather than chosen by the practitioner. In such scenarios, data could unavoidably be separated among workers due to limited memory capacity per worker or even because of data privacy issues. There, existing distributed methods will utterly fail due to dominant transfer costs across workers, or do not even apply. We propose a new approach to distributed fully connected neural network learning, called independent subnet training (IST), to handle these cases. In IST, the original network is decomposed into a set of narrow subnetworks with the same depth. These subnetworks are then trained locally before parameters are exchanged to produce new subnets and the training cycle repeats. Such a naturally "model parallel" approach limits memory usage by storing only a portion of network parameters on each device. Additionally, no requirements exist for sharing data between workers (i.e., subnet training is local and independent) and communication volume and frequency are reduced by decomposing the original network into independent subnets. These properties of IST can cope with issues due to distributed data, slow interconnects, or limited device memory, making IST a suitable approach for cases of mandatory distribution. We show experimentally that IST results in training times that are much lower than common distributed learning approaches.