Self-Tuned Descriptive Document Clustering Using a Predictive Network

Self-Tuned Descriptive Document Clustering Using a Predictive Network
复制标题

DOI:
10.1109/tkde.2017.2781721
复制
发表时间:
2018-10-01
影响因子:
8.9
通讯作者:
Goulermas, John Y.
Goulermas, John Y.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Brockmeier, Austin J.;Mu, Tingting;Goulermas, John Y.

文献摘要

被引文献

相似文献

描述性聚类包括自动将数据实例组织成聚类,并为每个聚类生成描述性摘要。描述应告知用户每个聚类的内容,而无需进一步检查具体实例,从而使用户能够快速扫描相关聚类。描述的选择通常依赖于启发式标准。我们将描述性聚类建模为自动编码器网络,该网络从聚类分配预测特征,并从特征子集预测聚类分配。用于预测聚类的特征子集用作其描述。对于文本文档,单词、短语或其他属性的出现或计数提供了具有可解释的特征标签的稀疏特征表示。在所提出的网络中,使用逻辑回归模型进行聚类预测,并且特征预测依赖于逻辑或多项式回归模型。优化这些模型会导致一个完全自调优的描述性聚类方法,自动选择聚类的数量和每个聚类的特征数量。我们将该方法应用于各种短文本文档,并表明所选择的聚类,所选择的特征子集证明,与一个有意义的主题组织。
Descriptive clustering consists of automatically organizing data instances into clusters and generating a descriptive summary for each cluster. The description should inform a user about the contents of each cluster without further examination of the specific instances, enabling a user to rapidly scan for relevant clusters. Selection of descriptions often relies on heuristic criteria. We model descriptive clustering as an auto-encoder network that predicts features from cluster assignments and predicts cluster assignments from a subset of features. The subset of features used for predicting a cluster serves as its description. For text documents, the occurrence or count of words, phrases, or other attributes provides a sparse feature representation with interpretable feature labels. In the proposed network, cluster predictions are made using logistic regression models, and feature predictions rely on logistic or multinomial regression models. Optimizing these models leads to a completely self-tuned descriptive clustering approach that automatically selects the number of clusters and the number of features for each cluster. We applied the methodology to a variety of short text documents and showed that the selected clustering, as evidenced by the selected feature subsets, are associated with a meaningful topical organization.