Symmetric Correspondence Topic Models for Multilingual Text Analysis

Symmetric Correspondence Topic Models for Multilingual Text Analysis
复制标题

DOI:
--
复制
发表时间:
2012-12
期刊:
--
影响因子:
--
通讯作者:
Kosuke Fukumasu;K. Eguchi;E. Xing
Kosuke Fukumasu;K. Eguchi;E. Xing
中科院分区:
其他
文献类型:
--
作者:
Kosuke Fukumasu;K. Eguchi;E. Xing

文献摘要

被引文献

相似文献

主题建模是一种广泛用于分析大型文本集合的方法。最近已经探索了少量多语言主题模型,以发现平行或可比文档中的潜在主题,例如维基百科。最初为结构化数据提出的其他主题模型也适用于多语言文档。对应潜狄利克雷分配(corlda)就是这样一个模型;但是,它需要事先指定一种枢轴语言。在CorrLDA的基础上,提出了一种新的主题模型对称对应LDA (SymCorrLDA),该模型引入了一个隐藏变量来控制枢轴语言。我们对从维基百科中提取的两个多语言可比较数据集进行了实验,并证明了SymCorrLDA比其他一些现有的多语言主题模型更有效。
Topic modeling is a widely used approach to analyzing large text collections. A small number of multilingual topic models have recently been explored to discover latent topics among parallel or comparable documents, such as in Wikipedia. Other topic models that were originally proposed for structured data are also applicable to multilingual documents. Correspondence Latent Dirichlet Allocation (CorrLDA) is one such model; however, it requires a pivot language to be specified in advance. We propose a new topic model, Symmetric Correspondence LDA (SymCorrLDA), that incorporates a hidden variable to control a pivot language, in an extension of CorrLDA. We experimented with two multilingual comparable datasets extracted from Wikipedia and demonstrate that SymCorrLDA is more effective than some other existing multilingual topic models.