Algorithms for Generalized Topic Modeling

Algorithms for Generalized Topic Modeling
复制标题

DOI:
10.1609/aaai.v32i1.11825
复制
发表时间:
2018-04
期刊:
--
影响因子:
--
通讯作者:
Avrim Blum;Nika Haghtalab
Avrim Blum;Nika Haghtalab
中科院分区:
其他
文献类型:
--
作者:
Avrim Blum;Nika Haghtalab

文献摘要

相似文献

最近有显着的活动,在开发算法与可证明的保证主题建模。在这项工作中,我们考虑了传统的主题建模框架,我们不再假设,词画i.i.d.的广泛推广。而是将主题视为段落序列上的复杂分布。由于人们甚至不能希望在一般情况下表示这样的分布(即使段落是使用一些自然特征表示的),我们的目标是直接学习一个预测器,该预测器在给定新文档的情况下准确预测其主题混合,而无需显式学习分布。我们提出了几个自然的条件下,可以做到这一点,从未标记的数据,并给出有效的算法,这样做,还讨论了问题,如噪声容限和样本的复杂性。更一般地说,我们的模型可以被视为机器学习中多视图或联合训练设置的泛化。
Recently there has been significant activity in developing algorithms with provable guarantees for topic modeling. In this work we consider a broad generalization of the traditional topic modeling framework, where we no longer assume that words are drawn i.i.d. and instead view a topic as a complex distribution over sequences of paragraphs. Since one could not hope to even represent such a distribution in general (even if paragraphs are given using some natural feature representation), we aim instead to directly learn a predictor that given a new document, accurately predicts its topic mixture, without learning the distributions explicitly. We present several natural conditions under which one can do this from unlabeled data only, and give efficient algorithms to do so, also discussing issues such as noise tolerance and sample complexity. More generally, our model can be viewed as a generalization of the multi-view or co-training setting in machine learning.