Federated Optimization: Distributed Machine Learning for On-Device Intelligence

Federated Optimization: Distributed Machine Learning for On-Device Intelligence
复制标题

DOI:
--
复制
发表时间:
2016-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Jakub Konecný;H. B. McMahan;Daniel Ramage;Peter Richtárik
Jakub Konecný;H. B. McMahan;Daniel Ramage;Peter Richtárik
中科院分区:
其他
文献类型:
--
作者:
Jakub Konecný;H. B. McMahan;Daniel Ramage;Peter Richtárik

文献摘要

被引文献

相似文献

我们引入了一种机器学习中分布式优化的新的且日益相关的设定,在这种设定中,定义优化的数据在极大量的节点上不均匀分布。目标是训练一个高质量的集中式模型。我们将这种设定称为联邦优化。在这种设定下,通信效率至关重要,并且最小化通信轮数是主要目标。当我们将训练数据保留在用户的移动设备本地,而不是将其记录到数据中心进行训练时,就会出现一个激励性的例子。在联邦优化中,设备被用作计算节点,对其本地数据进行计算以更新全局模型。我们假设网络中有极大量的设备——与给定服务的用户数量一样多,每个设备仅拥有可用总数据的极小一部分。特别是,我们预计本地可用的数据点数量比设备数量少得多。此外,由于不同用户以不同模式生成数据,因此可以合理地假设没有设备具有整体分布的代表性样本。我们表明现有算法不适用于这种设定,并提出了一种新算法,该算法在稀疏凸问题上显示出令人鼓舞的实验结果。这项工作也为联邦优化背景下所需的未来研究指明了方向。
We introduce a new and increasingly relevant setting for distributed optimization in machine learning, where the data defining the optimization are unevenly distributed over an extremely large number of nodes. The goal is to train a high-quality centralized model. We refer to this setting as Federated Optimization. In this setting, communication efficiency is of the utmost importance and minimizing the number of rounds of communication is the principal goal. A motivating example arises when we keep the training data locally on users' mobile devices instead of logging it to a data center for training. In federated optimziation, the devices are used as compute nodes performing computation on their local data in order to update a global model. We suppose that we have extremely large number of devices in the network --- as many as the number of users of a given service, each of which has only a tiny fraction of the total data available. In particular, we expect the number of data points available locally to be much smaller than the number of devices. Additionally, since different users generate data with different patterns, it is reasonable to assume that no device has a representative sample of the overall distribution. We show that existing algorithms are not suitable for this setting, and propose a new algorithm which shows encouraging experimental results for sparse convex problems. This work also sets a path for future research needed in the context of \federated optimization.