Distributed Semi-Supervised Learning With Consensus Consistency on Edge Devices

Distributed Semi-Supervised Learning With Consensus Consistency on Edge Devices
复制标题

DOI:
10.1109/tpds.2023.3340707
复制
发表时间:
2024-02
影响因子:
5.3
通讯作者:
Hao-Rui Chen;Lei Yang;Xinglin Zhang;Jiaxing Shen;Jiannong Cao
Hao-Rui Chen;Lei Yang;Xinglin Zhang;Jiaxing Shen;Jiannong Cao
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hao-Rui Chen;Lei Yang;Xinglin Zhang;Jiaxing Shen;Jiannong Cao

文献摘要

相似文献

分布式学习在边缘计算领域得到了越来越多的研究,它使边缘设备能够在不交换私有数据的情况下协同学习模型。然而,现有方法假设边缘设备拥有的私有数据都被标记,而现实情况是大量私有数据未被标记并被利用,这导致性能不理想。为了克服这一限制,我们研究了一个新的实际问题,分布式半监督学习(DSSL),在每个设备上使用混合的私有标记和未标记数据协同学习模型。我们还提出了一种新的方法DistMatch,该方法利用每个设备上的私有未标记数据,在邻近设备的模型的帮助下进行自我训练。DistMatch通过正确地平均这些接收到的模型的预测,为未标记的数据生成伪标签。此外,为了避免使用错误的伪标签进行自我训练,DistMatch提出了共识一致性损失来过滤一致性高的伪标签,并强制训练模型的输出与这些伪标签一致。通过我们自己开发的测试平台进行的广泛评估结果表明,该方法在常用的图像分类基准数据集上优于所有基线。
Distributed learning has been increasingly studied in edge computing, enabling edge devices to learn a model collaboratively without exchanging their private data. However, existing approaches assume the private data owned by edge devices are all labeled while the reality is that massive private data are unlabeled and remain to be utilized, which leads to suboptimal performance. To overcome this limitation, we study a new practical problem, Distributed Semi-Supervised Learning (DSSL), to learn models collaboratively with mixed private labeled and unlabeled data on each device. We also propose a novel method DistMatch that exploits private unlabeled data by self-training on each device with the help of models from neighboring devices. DistMatch generates pseudo-labels for unlabeled data by properly averaging the predictions of these received models. Furthermore, to avoid self-training with wrong pseudo-labels, DistMatch proposes a consensus consistency loss to filter pseudo-labels with high consensus and force the output of the trained model to be consistent with these pseudo-labels. Extensive evaluation results via our self-developed testbed indicate the proposed method outperforms all baselines on commonly used image classification benchmark datasets.