Distributed Minimum Error Entropy Algorithms

Distributed Minimum Error Entropy Algorithms
复制标题

分布式最小误差熵算法

DOI:
--
复制
发表时间:
2020
影响因子:
6
通讯作者:
Qiang Wu
Qiang Wu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Xin Guo;Ting Hu;Qiang Wu

文献摘要

相似文献

最小误差熵原理是信息理论学习中的一个重要方法。由于其对噪声的鲁棒性,在各个领域得到了广泛的应用和研究。在本文中,我们研究了一种基于再生核的分布式并行计算算法,DSPRING,which.is,设计用于全监督数据和半监督数据。采用分治的方法,所以没有节点间的通信开销。与其他分布式算法类似,DMEE显着降低了单个计算节点的计算复杂度和内存需求。对于完全监督的数据,我们证明的学习率等于经典的逐点基于核的回归的极大极小最优学习率。在半监督学习场景下,我们证明了DSTO有效地利用了未标记数据,在弱规则性假设的情况下,额外的未标记数据显著提高了DSTO的学习率。其次,有了足够的未标记数据,labeled.data可以被分发到更多的计算节点,每个节点只需要O(1)个标签,而不会破坏标签数量方面的学习率。这一结论克服了未标记数据量饱和的现象。它与正则化最小二乘的最近结果(Lin和Zhou,2018)相似,并表明未标记数据的信息化是解决去中心化数据源隐私保护问题的一种方法。我们的工作涉及成对学习和非凸损失。理论分析是通过积分算子的分布U-统计量和误差分解技术实现的。关键词:信息论学习,最小误差熵,分布式方法,半监督数据,再生核希尔伯特空间
Minimum Error Entropy (MEE) principle is an important approach in Information Theoretical.Learning (ITL). It is widely applied and studied in various elds for its robustness to noise..In this paper, we study a reproducing kernel-based distributed MEE algorithm, DMEE, which.is designed to work with both fully supervised data and semi-supervised data. The divide-and-.conquer approach is employed, so there is no inter-node communication overhead. Similar as other.distributed algorithms, DMEE signicantly reduces the computational complexity and memory.requirement on single computing nodes. With fully supervised data, our proved learning rates.equal the minimax optimal learning rates of the classical pointwise kernel-based regressions. Under.the semi-supervised learning scenarios, we show that DMEE exploits unlabeled data eectively, in.the sense that rst, under the settings with weak regularity assumptions, additional unlabeled data.signicantly improves the learning rates of DMEE. Second, with sucient unlabeled data, labeled.data can be distributed to many more computing nodes, that each node takes only O(1) labels,.without spoiling the learning rates in terms of the number of labels. This conclusion overcomes.the saturation phenomenon in unlabeled data size. It parallels a recent results for regularized least.squares (Lin and Zhou, 2018), and suggests that an in.ation of unlabeled data is a solution to.the MEE learning problems with decentralized data source for the concerns of privacy protection..Our work refers to pairwise learning and non-convex loss. The theoretical analysis is achieved by.distributed U-statistics and error decomposition techniques in integral operators..Keywords: Information theoretic learning, minimum error entropy, distributed method, semi-.supervised data, reproducing kernel Hilbert space