Coresets for Archetypal Analysis

Coresets for Archetypal Analysis
复制标题

用于原型分析的核心集

DOI:
--
复制
发表时间:
2019
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
Ulf Brefeld
Ulf Brefeld
中科院分区:
--
文献类型:
--
作者:
Sebastian Mair;Ulf Brefeld

文献摘要

被引文献

相似文献

原型分析将实例表示为位于数据凸包边界上的原型(原型)的线性混合。因此,原型通常比其他矩阵分解技术计算的因子更容易解释。然而,由于额外的凸性保持约束,可解释性伴随着较高的计算成本。在本文中,我们提出了用于原型分析的有效核心集。通过显示 k 均值上限原型分析的量化误差得出理论保证;可证明的绝对核心集的计算只需对数据进行两次传递即可执行。根据经验,我们表明核心集可以提高多个数据集的性能。
Archetypal analysis represents instances as linear mixtures of prototypes (the archetypes) that lie on the boundary of the convex hull of the data. Archetypes are thus often better interpretable than factors computed by other matrix factorization techniques. However, the interpretability comes with high computational cost due to additional convexity-preserving constraints. In this paper, we propose efficient coresets for archetypal analysis. Theoretical guarantees are derived by showing that quantization errors of k-means upper bound archetypal analysis; the computation of a provable absolute-coreset can be performed in only two passes over the data. Empirically, we show that the coresets lead to improved performance on several data sets.