Leveraging Spatial and Temporal Correlations in Sparsified Mean Estimation

Leveraging Spatial and Temporal Correlations in Sparsified Mean Estimation
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Divyansh Jhunjhunwala;Ankur Mallick;Advait Gadhikar;S. Kadhe;Gauri Joshi
Divyansh Jhunjhunwala;Ankur Mallick;Advait Gadhikar;S. Kadhe;Gauri Joshi
中科院分区:
其他
文献类型:
--
作者:
Divyansh Jhunjhunwala;Ankur Mallick;Advait Gadhikar;S. Kadhe;Gauri Joshi

文献摘要

相似文献

我们研究了在中央服务器上估计分布在几个节点上的一组向量(每个节点一个向量)的平均值的问题。当向量是高维的时候,发送整个向量的通信成本可能是令人望而却步的,并且它们可能必须使用稀疏化技术。虽然大多数现有的稀疏均值估计工作与数据向量的特性无关,但在许多实际应用中,如联合学习,数据向量中可能存在空间相关性(不同节点发送的向量的相似性)或时间相关性(单个节点在算法不同迭代期间发送的数据的相似性)。我们只需修改服务器使用的解码方法来估计平均值,即可利用这些相关性。我们分析了由此产生的估计误差,并对PCA、K-均值和Logistic回归进行了实验,结果表明,我们的估计器的性能始终优于更复杂和更昂贵的稀疏方法。
We study the problem of estimating at a central server the mean of a set of vectors distributed across several nodes (one vector per node). When the vectors are high-dimensional, the communication cost of sending entire vectors may be prohibitive, and it may be imperative for them to use sparsification techniques. While most existing work on sparsified mean estimation is agnostic to the characteristics of the data vectors, in many practical applications such as federated learning, there may be spatial correlations (similarities in the vectors sent by different nodes) or temporal correlations (similarities in the data sent by a single node over different iterations of the algorithm) in the data vectors. We leverage these correlations by simply modifying the decoding method used by the server to estimate the mean. We provide an analysis of the resulting estimation error as well as experiments for PCA, K-Means and Logistic Regression, which show that our estimators consistently outperform more sophisticated and expensive sparsification methods.