Localized user-driven topic discovery via boosted ensemble of nonnegative matrix factorization

Localized user-driven topic discovery via boosted ensemble of nonnegative matrix factorization
复制标题

DOI:
10.1007/s10115-017-1147-9
复制
发表时间:
2018-09-01
影响因子:
2.7
通讯作者:
Choo, Jaegul
Choo, Jaegul
中科院分区:
计算机科学4区
文献类型:
--
作者:
Suh, Sangho;Shin, Sungbok;Choo, Jaegul

文献摘要

被引文献

相似文献

非负矩阵分解(NMF)被广泛地应用于大规模文档语料库的主题建模中,从NMF中提取一组潜在主题的低秩因式矩阵。然而,由此产生的主题通常只传达关于文档的一般性的冗余信息,而不是可能对用户有潜在意义的次要信息。为了解决这个问题,我们提出了一种新的基于非负矩阵分解的集成方法,它可以发现有意义的局部主题。我们的方法利用集成模型的思想,这在有监督学习中显示了优势,进入了无监督主题建模环境。也就是说,我们的模型在给定从前面阶段获得的残差矩阵的情况下连续执行NMF,并生成主题集序列。我们采用的更新算法在两个方面都是新颖的。第一个是利用最先进的梯度提升模型启发的残差矩阵,第二个是在给定的矩阵上应用复杂的局部加权方案来增强主题的局部性,这反过来又提供用户感兴趣的高质量、集中的主题。随后,我们通过添加基于关键字和文档的用户交互来扩展这一集成模型,以引入用户驱动的主题发现。
Nonnegative matrix factorization (NMF) has been widely used in topic modeling of large-scale document corpora, where a set of underlying topics are extracted by a low-rank factor matrix from NMF. However, the resulting topics often convey only general, thus redundant information about the documents rather than information that might be minor, but potentially meaningful to users. To address this problem, we present a novel ensemble method based on nonnegative matrix factorization that discovers meaningful local topics. Our method leverages the idea of an ensemble model, which has shown advantages in supervised learning, into an unsupervised topic modeling context. That is, our model successively performs NMF given a residual matrix obtained from previous stages and generates a sequence of topic sets. The algorithm we employ to update is novel in two aspects. The first lies in utilizing the residual matrix inspired by a state-of-the-art gradient boosting model, and the second stems from applying a sophisticated local weighting scheme on the given matrix to enhance the locality of topics, which in turn delivers high-quality, focused topics of interest to users. We subsequently extend this ensemble model by adding keyword- and document-based user interaction to introduce user-driven topic discovery.