Explainable k-means: don’t be greedy, plant bigger trees!

Explainable k-means: don’t be greedy, plant bigger trees!
复制标题

DOI:
10.1145/3519935.3520056
复制
发表时间:
2021-11
期刊:
Proceedings of the 54th Annual ACM SIGACT Symposium on Theory of Computing
影响因子:
--
通讯作者:
K. Makarychev;Liren Shan
K. Makarychev;Liren Shan
中科院分区:
其他
文献类型:
--
作者:
K. Makarychev;Liren Shan

文献摘要

被引文献

相似文献

我们提供了一个新的双重标准(Log2 K)竞争性算法,以解释可解释的K-Means。并了解(阈值)决策树或图表。该群集的最佳非双重标准算法,用于竞争性的可解释的聚类(K) (其中δ∈(0,1)是算法的参数)。绑定几乎是最佳的。
We provide a new bi-criteria Õ(log2 k) competitive algorithm for explainable k-means clustering. Explainable k-means was recently introduced by Dasgupta, Frost, Moshkovitz, and Rashtchian (ICML 2020). It is described by an easy to interpret and understand (threshold) decision tree or diagram. The cost of the explainable k-means clustering equals to the sum of costs of its clusters; and the cost of each cluster equals the sum of squared distances from the points in the cluster to the center of that cluster. The best non bi-criteria algorithm for explainable clustering Õ(k) competitive, and this bound is tight. Our randomized bi-criteria algorithm constructs a threshold decision tree that partitions the data set into (1+δ)k clusters (where δ∈ (0,1) is a parameter of the algorithm). The cost of this clustering is at most Õ(1/δ· log2 k) times the cost of the optimal unconstrained k-means clustering. We show that this bound is almost optimal.