Sequential Modeling of Topic Dynamics with Multiple Timescales

Sequential Modeling of Topic Dynamics with Multiple Timescales
复制标题

DOI:
10.1145/2086737.2086739
复制
发表时间:
2012-02
期刊:
ACM Trans. Knowl. Discov. Data
影响因子:
--
通讯作者:
Tomoharu Iwata;Takeshi Yamada;Yasushi Sakurai;N. Ueda
Tomoharu Iwata;Takeshi Yamada;Yasushi Sakurai;N. Ueda
中科院分区:
其他
文献类型:
--
作者:
Tomoharu Iwata;Takeshi Yamada;Yasushi Sakurai;N. Ueda

文献摘要

被引文献

相似文献

我们提出了一个在线主题模型,顺序分析的时间演变的主题在文档集合。主题自然会随着多个时间尺度而演变。例如,有些词可能在一百年内一直使用,而另一些词则在几天内出现和消失。因此,在所提出的模型中,当前特定于主题的分布在单词被假定为基于前一个时代的多尺度单词分布而生成。考虑到长期和短期的依赖性产生一个更强大的模型。我们推导出有效的在线推理过程的基础上的随机EM算法,在该算法中,该模型是顺序更新使用新获得的数据,这意味着过去的数据不需要作出推断。我们证明了所提出的方法的预测性能和计算效率方面的有效性,通过检查收集的真实的文件的时间戳。
We propose an online topic model for sequentially analyzing the time evolution of topics in document collections. Topics naturally evolve with multiple timescales. For example, some words may be used consistently over one hundred years, while other words emerge and disappear over periods of a few days. Thus, in the proposed model, current topic-specific distributions over words are assumed to be generated based on the multiscale word distributions of the previous epoch. Considering both the long- and short-timescale dependency yields a more robust model. We derive efficient online inference procedures based on a stochastic EM algorithm, in which the model is sequentially updated using newly obtained data; this means that past data are not required to make the inference. We demonstrate the effectiveness of the proposed method in terms of predictive performance and computational efficiency by examining collections of real documents with timestamps.