A network approach to topic models.

A network approach to topic models.
复制标题

DOI:
10.1126/sciadv.aaq1360
复制
发表时间:
2018-07
期刊:
影响因子:
13.6
通讯作者:
Altmann EG
Altmann EG
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Gerlach M;Peixoto TP;Altmann EG

文献摘要

参考文献

被引文献

相似文献

一种新的主题模型方法通过词文档网络中的社区检测来发现主题。从非结构化文本中提取有用的信息是现代计算和科学的主要挑战之一。主题模型是一种流行的机器学习方法,它可以推断文档集合的潜在主题结构。尽管他们的成功,特别是最广泛使用的变量称为潜在狄利克雷分配(LDA)和众多的应用在社会学,历史学和语言学,主题模型是已知的遭受严重的概念和实际问题,例如,缺乏理由的贝叶斯先验,差异与统计特性的真实的文本,以及无法正确地选择主题的数量。我们获得了一个新的观点,确定主题结构的问题,通过将其与在复杂网络中找到社区的问题。我们通过将文本语料库表示为文档和单词的二分网络来实现这一点。通过调整现有的社区检测方法(使用随机块模型(SBM)与非参数先验),我们获得了一个更通用和原则性的框架主题建模(例如,它自动检测主题的数量和分层聚类的话和文件)。人工和真实的语料库的分析表明,我们的SBM方法导致更好的主题模型比LDA的统计模型选择。我们的工作展示了如何将社区检测和主题建模的方法正式联系起来,从而打开了这两个领域之间相互促进的可能性。
A new approach to topic models finds topics through community detection in word-document networks. One of the main computational and scientific challenges in the modern age is to extract useful information from unstructured texts. Topic models are one popular machine-learning approach that infers the latent topical structure of a collection of documents. Despite their success—particularly of the most widely used variant called latent Dirichlet allocation (LDA)—and numerous applications in sociology, history, and linguistics, topic models are known to suffer from severe conceptual and practical problems, for example, a lack of justification for the Bayesian priors, discrepancies with statistical properties of real texts, and the inability to properly choose the number of topics. We obtain a fresh view of the problem of identifying topical structures by relating it to the problem of finding communities in complex networks. We achieve this by representing text corpora as bipartite networks of documents and words. By adapting existing community-detection methods (using a stochastic block model (SBM) with nonparametric priors), we obtain a more versatile and principled framework for topic modeling (for example, it automatically detects the number of topics and hierarchically clusters both the words and documents). The analysis of artificial and real corpora demonstrates that our SBM approach leads to better topic models than LDA in terms of statistical model selection. Our work shows how to formally relate methods from community detection and topic modeling, opening the possibility of cross-fertilization between these two fields.
DOI: 10.1073/pnas.0307752101
发表时间: 2004-04-06
影响因子: 11.1
作者:
Griffiths, TL;Steyvers, M
通讯作者: Steyvers, M
DOI: 10.1103/physreve.97.052303
发表时间: 2018-05-10
期刊: PHYSICAL REVIEW E
影响因子: 2.4
作者:
Courtney, Owen T.;Bianconi, Ginestra
通讯作者: Bianconi, Ginestra
DOI: 10.1162/jmlr.2003.3.4-5.993
发表时间: 2003-05-15
影响因子: 6
作者:
Blei, DM;Ng, AY;Jordan, MI
通讯作者: Jordan, MI
DOI: 10.1103/physreve.70.025101
发表时间: 2004-08-01
期刊: PHYSICAL REVIEW E
影响因子: 2.4
作者:
Guimerà, R;Sales-Pardo, M;Amaral, LAN
通讯作者: Amaral, LAN
DOI: 10.1103/physreve.84.036103
发表时间: 2011-09-08
期刊: PHYSICAL REVIEW E
影响因子: 2.4
作者:
Ball, Brian;Karrer, Brian;Newman, M. E. J.
通讯作者: Newman, M. E. J.