Cloud K-SVD: A Collaborative Dictionary Learning Algorithm for Big, Distributed Data

Cloud K-SVD: A Collaborative Dictionary Learning Algorithm for Big, Distributed Data
复制标题

DOI:
10.1109/tsp.2015.2472372
复制
发表时间:
2014-12
影响因子:
5.4
通讯作者:
Haroon Raja;W. Bajwa
Haroon Raja;W. Bajwa
中科院分区:
工程技术1区
文献类型:
--
作者:
Haroon Raja;W. Bajwa

文献摘要

被引文献

相似文献

本文研究了分布式大数据的数据自适应表示问题。假设一些地理上分布的、相互连接的站点具有大量的本地数据,并且它们有兴趣协作地学习这些数据的低维几何结构。与以前的作品基于子空间的数据表示,本文重点讨论的几何结构的子空间(UoS)的联盟。在这方面,它提出了一种分布式算法云K-SVD-UoS结构的基础上的分布式数据的兴趣协作学习。云K-SVD的目标是在每个单独的站点学习一个公共的过完备字典,使得分布式数据中的每个样本都可以通过学习字典的少量原子来表示。Cloud K-SVD实现了这一目标,而无需在研究中心之间交换单个样本。这使得它适用于由于隐私问题或大量数据而不鼓励共享原始数据的应用程序。本文还提供了云K-SVD的分析,深入了解其属性以及在各个站点从集中式解决方案中学习的字典的偏差,这些偏差是根据本地/全局数据和互连拓扑的不同度量。最后,本文数值说明了云K-SVD对真实的和合成的分布式数据的有效性。
This paper studies the problem of data-adaptive representations for big, distributed data. It is assumed that a number of geographically-distributed, interconnected sites have massive local data and they are interested in collaboratively learning a low-dimensional geometric structure underlying these data. In contrast with previous works on subspace-based data representations, this paper focuses on the geometric structure of a union of subspaces (UoS). In this regard, it proposes a distributed algorithm-termed cloud K-SVD-for collaborative learning of a UoS structure underlying distributed data of interest. The goal of cloud K-SVD is to learn a common overcomplete dictionary at each individual site such that every sample in the distributed data can be represented through a small number of atoms of the learned dictionary. Cloud K-SVD accomplishes this goal without requiring exchange of individual samples between sites. This makes it suitable for applications where sharing of raw data is discouraged due to either privacy concerns or large volumes of data. This paper also provides an analysis of cloud K-SVD that gives insights into its properties as well as deviations of the dictionaries learned at individual sites from a centralized solution in terms of different measures of local/global data and topology of interconnections. Finally, the paper numerically illustrates the efficacy of cloud K-SVD on real and synthetic distributed data.