Label consistency in overfitted generalized $k$-means

Label consistency in overfitted generalized $k$-means
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Linfan Zhang;A. Amini
Linfan Zhang;A. Amini
中科院分区:
其他
文献类型:
--
作者:
Linfan Zhang;A. Amini

文献摘要

相似文献

我们为广义k-均值问题的标签一致性提供了理论上的保证,重点讨论了算法使用的簇数大于基本事实的OverfiDisted情况。我们给出了估计的标号接近真实簇标号的条件。我们同时考虑了标号的精确恢复和近似恢复。我们的结果对k-均值问题的任何常数因子近似都成立。结果也是无模型的,并且只基于数据点到真实集群中心的最大或平均距离的界限。这些中心本身是松散的defiNed,并且可以被认为是前述距离可以控制的任何一组点。通过对一些流形聚类问题的应用,证明了结果的有效性
We provide theoretical guarantees for label consistency in generalized k -means problems, with an emphasis on the overfitted case where the number of clusters used by the algorithm is more than the ground truth. We provide conditions under which the estimated labels are close to a refinement of the true cluster labels. We consider both exact and approximate recovery of the labels. Our results hold for any constant-factor approximation to the k -means problem. The results are also model-free and only based on bounds on the maximum or average distance of the data points to the true cluster centers. These centers themselves are loosely defined and can be taken to be any set of points for which the aforementioned distances can be controlled. We show the usefulness of the results with applications to some manifold clustering problems