Deep Descriptive Clustering

Deep Descriptive Clustering
复制标题

DOI:
10.24963/ijcai.2021/460
复制
发表时间:
2021-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Hongjing Zhang;I. Davidson
Hongjing Zhang;I. Davidson
中科院分区:
其他
文献类型:
--
作者:
Hongjing Zhang;I. Davidson

文献摘要

相似文献

最近的工作可解释的聚类允许描述集群时的功能是可解释的。然而,许多现代机器学习都集中在复杂的数据上,如图像、文本和图形,这些数据使用了深度学习,但数据的原始特征是不可解释的。本文探讨了一种新的设置复杂的数据进行聚类,同时使用可解释的标签生成解释。我们提出了深度描述性聚类,它在复杂数据上执行子符号表示学习,同时基于符号数据生成解释。我们形成良好的集群,通过最大化的经验分布之间的互信息的输入和聚类目标的诱导聚类标签。我们通过求解一个整数线性规划来生成解释,该规划为每个聚类生成简洁和正交的描述。最后,我们允许的解释,通知更好的聚类提出了一种新的成对损失与自我生成的约束,以最大限度地提高聚类和解释模块的一致性。在公开数据上的实验结果表明,我们的模型在聚类性能上优于竞争基线,同时提供高质量的聚类级解释。
Recent work on explainable clustering allows describing clusters when the features are interpretable. However, much modern machine learning focuses on complex data such as images, text, and graphs where deep learning is used but the raw features of data are not interpretable. This paper explores a novel setting for performing clustering on complex data while simultaneously generating explanations using interpretable tags. We propose deep descriptive clustering that performs sub-symbolic representation learning on complex data while generating explanations based on symbolic data. We form good clusters by maximizing the mutual information between empirical distribution on the inputs and the induced clustering labels for clustering objectives. We generate explanations by solving an integer linear programming that generates concise and orthogonal descriptions for each cluster. Finally, we allow the explanation to inform better clustering by proposing a novel pairwise loss with self-generated constraints to maximize the clustering and explanation module's consistency. Experimental results on public data demonstrate that our model outperforms competitive baselines in clustering performance while offering high-quality cluster-level explanations.