Diversifying Restricted Boltzmann Machine for Document Modeling

Diversifying Restricted Boltzmann Machine for Document Modeling
复制标题

DOI:
10.1145/2783258.2783264
复制
发表时间:
2015-08
期刊:
Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
P. Xie;Yuntian Deng;E. Xing
P. Xie;Yuntian Deng;E. Xing
中科院分区:
其他
文献类型:
--
作者:
P. Xie;Yuntian Deng;E. Xing

文献摘要

被引文献

相似文献

受限玻尔兹曼机(RBM)在文档建模中显示出极大的有效性。它利用隐藏单元来发现潜在的主题,并可以学习文档的紧凑语义表示,极大地方便了文档的检索、聚类和分类。文本语料库中主题的流行度(或频率)通常遵循幂律分布,其中少数主导主题非常频繁地出现,而大多数主题(在长尾区域)的概率很低。由于这种不平衡,RBM倾向于学习多个冗余隐藏单元来最好地代表主导主题,而忽略长尾区域的隐藏单元,这使得学习到的表示是冗余的和无信息的。为了解决这一问题,我们提出了多元化RBM (DRBM),将隐藏单元多样化,使其不仅覆盖主导话题,而且覆盖长尾区域的话题。我们定义了一个多样性度量,并将其用作正则化器,以鼓励隐藏单元具有多样性。由于多样性指标难以直接优化,我们转而优化其下界,并证明随着投影梯度上升最大化下界可以增加该多样性指标。文档检索和聚类实验表明,多样化可以大大提高DRBM的文档建模能力。
Restricted Boltzmann Machine (RBM) has shown great effectiveness in document modeling. It utilizes hidden units to discover the latent topics and can learn compact semantic representations for documents which greatly facilitate document retrieval, clustering and classification. The popularity (or frequency) of topics in text corpora usually follow a power-law distribution where a few dominant topics occur very frequently while most topics (in the long-tail region) have low probabilities. Due to this imbalance, RBM tends to learn multiple redundant hidden units to best represent dominant topics and ignore those in the long-tail region, which renders the learned representations to be redundant and non-informative. To solve this problem, we propose Diversified RBM (DRBM) which diversifies the hidden units, to make them cover not only the dominant topics, but also those in the long-tail region. We define a diversity metric and use it as a regularizer to encourage the hidden units to be diverse. Since the diversity metric is hard to optimize directly, we instead optimize its lower bound and prove that maximizing the lower bound with projected gradient ascent can increase this diversity metric. Experiments on document retrieval and clustering demonstrate that with diversification, the document modeling power of DRBM can be greatly improved.