Truthful Incentive Mechanism for Federated Learning with Crowdsourced Data Labeling

Truthful Incentive Mechanism for Federated Learning with Crowdsourced Data Labeling
复制标题

DOI:
10.1109/infocom53939.2023.10228923
复制
发表时间:
2023-01
期刊:
IEEE INFOCOM 2023 - IEEE Conference on Computer Communications
影响因子:
--
通讯作者:
Yuxi Zhao;Xiaowen Gong;S. Mao
Yuxi Zhao;Xiaowen Gong;S. Mao
中科院分区:
其他
文献类型:
--
作者:
Yuxi Zhao;Xiaowen Gong;S. Mao

文献摘要

相似文献

联邦学习 (FL) 最近成为一种有前途的范例,它以分布式方式在客户端设备上训练机器学习 (ML) 模型,而无需将客户端数据传输到 FL 服务器。在机器学习的许多应用中(例如图像分类),训练数据的标签需要由人类代理手动生成(例如识别和注释图像中的对象),这通常成本高昂且容易出错。在本文中,我们通过众包数据标记来研究 FL,其中每个参与 FL 的客户端的本地数据由客户端手动标记。我们考虑客户的策略行为,他们可能没有在本地数据标记和本地模型计算(通过随机梯度计算中使用的小批量大小进行量化)方面做出预期的努力,并且可能会向 FL 服务器错误报告其本地模型。我们首先将训练损失的性能界限描述为客户数据标记工作、本地计算工作和报告的本地模型的函数,揭示了这些因素对训练损失的影响。有了这些见解,我们设计了标签和计算工作以及本地模型启发(LCEME)机制,激励战略客户在本地数据标签和本地模型计算方面按照服务器的要求做出真实的努力,并向服务器报告真实的本地模型。 LCEME 机制的真实设计利用了训练损失对客户隐藏努力和私有本地模型的非平凡依赖性,并克服了客户努力和本地模型联合引发中的复杂耦合。在 LCEME 机制下,我们描述了服务器的最佳本地计算工作量分配并分析了它们的性能。我们使用众包数据标记和基于 MNIST 的手写数字分类的 LCEME 机制来评估所提出的 FL 算法。结果证实了所提出的方法提高了学习准确性和成本效益。
Federated learning (FL) has recently emerged as a promising paradigm that trains machine learning (ML) models on clients' devices in a distributed manner without the need of transmitting clients' data to the FL server. In many applications of ML (e.g., image classification), the labels of training data need to be generated manually by human agents (e.g., recognizing and annotating objects in an image), which are usually costly and error-prone. In this paper, we study FL with crowdsourced data labeling where the local data of each participating client of FL are labeled manually by the client. We consider the strategic behavior of clients who may not make desired effort in their local data labeling and local model computation (quantified by the mini-batch size used in the stochastic gradient computation), and may misreport their local models to the FL server. We first characterize the performance bounds on the training loss as a function of clients' data labeling effort, local computation effort, and reported local models, which reveal the impacts of these factors on the training loss. With these insights, we devise Labeling and Computation Effort and local Model Elicitation (LCEME) mechanisms which incentivize strategic clients to make truthful efforts as desired by the server in local data labeling and local model computation, and also report true local models to the server. The truthful design of the LCEME mechanism exploits the non-trivial dependence of the training loss on clients' hidden efforts and private local models, and overcomes the intricate coupling in the joint elicitation of clients' efforts and local models. Under the LCEME mechanism, we characterize the server’s optimal local computation effort assignments and analyze their performance. We evaluate the proposed FL algorithms with crowdsourced data labeling and the LCEME mechanism for the MNIST-based hand-written digit classification. The results corroborate the improved learning accuracy and cost-effectiveness of the proposed approaches.