An Evaluative Measure of Clustering Methods Incorporating Hyperparameter Sensitivity

An Evaluative Measure of Clustering Methods Incorporating Hyperparameter Sensitivity
复制标题

DOI:
10.1609/aaai.v36i7.20747
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Siddhartha Mishra;Nicholas Monath;Michael Boratko;Ari Kobren;A. McCallum
Siddhartha Mishra;Nicholas Monath;Michael Boratko;Ari Kobren;A. McCallum
中科院分区:
其他
文献类型:
--
作者:
Siddhartha Mishra;Nicholas Monath;Michael Boratko;Ari Kobren;A. McCallum

文献摘要

被引文献

相似文献

聚类算法通常使用与真实聚类分配进行比较的度量进行评估,例如兰德指数和NMI。然而,对于不同的超参数,算法性能可能会有很大的不同,因此基于这些指标的最佳性能的模型选择与这些算法在实践中的应用方式不一致,其中标签不可用,并且调优通常更像是艺术而不是科学。因此,我们希望不仅在优化性能方面对集群算法进行比较,而且还在实践中获得这种性能的现实程度方面进行比较。我们提出了一个评估的聚类方法捕捉这种易于调整的建模预期的最佳聚类得分在给定的计算预算。为了鼓励通过建议的度量标准,经典的聚类评估,我们提供了一个可扩展的基准框架。我们进行了广泛的实证评估,我们提出的度量流行的聚类算法在一个大的集合,从不同的领域的数据集,并观察到我们的新指标导致几个值得注意的观察。
Clustering algorithms are often evaluated using metrics which compare with ground-truth cluster assignments, such as Rand index and NMI. Algorithm performance may vary widely for different hyperparameters, however, and thus model selection based on optimal performance for these metrics is discordant with how these algorithms are applied in practice, where labels are unavailable and tuning is often more art than science. It is therefore desirable to compare clustering algorithms not only on their optimally tuned performance, but also some notion of how realistic it would be to obtain this performance in practice. We propose an evaluation of clustering methods capturing this ease-of-tuning by modeling the expected best clustering score under a given computation budget. To encourage the adoption of the proposed metric alongside classic clustering evaluations, we provide an extensible benchmarking framework. We perform an extensive empirical evaluation of our proposed metric on popular clustering algorithms over a large collection of datasets from different domains, and observe that our new metric leads to several noteworthy observations.