Random Projections for k-Means: Maintaining Coresets Beyond Merge & Reduce

Random Projections for k-Means: Maintaining Coresets Beyond Merge & Reduce
复制标题

k-Means 的随机投影:在合并之外维护核心集

DOI:
--
复制
发表时间:
2015
期刊:
arXiv.org
影响因子:
--
通讯作者:
Chris Schwiegelshohn
Chris Schwiegelshohn
中科院分区:
--
文献类型:
--
作者:
Marc Bury;Chris Schwiegelshohn

文献摘要

被引文献

相似文献

我们给出了一个新的建设一个小的空间摘要满足coreset保证的数据集的$k $-means目标函数。离线构建所需的点数以$为单位 ilde {O}(k ∈ ^{-2} min(d,k ∈ ^{-2}))$,它是所有可用构造中最小的。 除了两个结构与维度指数相关外,所有已知的coreset都通过merge和reduce框架在数据流中维护,这会导致对$log n $的大空间依赖。相反,我们的建设至关重要地依赖于约翰逊-林登施特劳斯类型的嵌入,结合在线算法的结果给我们一个新的技术,有效地维护coresets在数据流中,而不依赖于合并和减少。我们的算法在一个数据流中存储的最终点数是以$ ilde {O}(k^2 ^{-2} log^2 n min(d,k^2 ^{-2}))$.
We give a new construction for a small space summary satisfying the coreset guarantee of a data set with respect to the $k$-means objective function. The number of points required in an offline construction is in $ ilde{O}(k epsilon^{-2}min(d,kepsilon^{-2}))$ which is minimal among all available constructions. Aside from two constructions with exponential dependence on the dimension, all known coresets are maintained in data streams via the merge and reduce framework, which incurs are large space dependency on $log n$. Instead, our construction crucially relies on Johnson-Lindenstrauss type embeddings which combined with results from online algorithms give us a new technique for efficiently maintaining coresets in data streams without relying on merge and reduce. The final number of points stored by our algorithm in a data stream is in $ ilde{O}(k^2 epsilon^{-2} log^2 n min(d,kepsilon^{-2}))$.