Visualization Techniques for Topic Model Checking

Visualization Techniques for Topic Model Checking
复制标题

主题模型检查的可视化技术

DOI:
10.1609/aaai.v29i1.9268
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
C. Allen
C. Allen
中科院分区:
--
文献类型:
--
作者:
J. Murdock;C. Allen

文献摘要

被引文献

相似文献

在许多方面,主题模型对于建模者和最终用户来说仍然是一个黑匣子。从建模者的角度来看,必须做出许多缺乏明确理由且相互作用不明确的决策 - 例如,算法应该找到多少个主题(K),要忽略哪些单词(又名“停止列表”),以及是否足以运行一次或多次建模过程,由于近似贝叶斯先验的算法而产生不同的结果。此外,不同参数设置的结果难以分析、总结和可视化,使得模型比较变得困难。从最终用户的角度来看,很难理解为什么模型会表现出这样的效果,并且信息论相似性度量并不完全符合主题的人文解释。我们推出了主题浏览器,它推进了文档-文档和主题-文档关系主题模型可视化的最先进技术。它以一种促进对语料库和模型的深入理解的方式将主题模型带入生活,允许用户生成解释性假设并建议进一步的实验。这些工具是评估主题建模是否适合人工智能和认知建模应用的技术的重要一步。
Topic models remain a black box both for modelers and for end users in many respects. From the modelers' perspective, many decisions must be made which lack clear rationales and whose interactions are unclear — for example, how many topics the algorithms should find (K), which words to ignore (aka the "stop list"), and whether it is adequate to run the modeling process once or multiple times, producing different results due to the algorithms that approximate the Bayesian priors. Furthermore, the results of different parameter settings are hard to analyze, summarize, and visualize, making model comparison difficult. From the end users' perspective, it is hard to understand why the models perform as they do, and information-theoretic similarity measures do not fully align with humanistic interpretation of the topics. We present the Topic Explorer, which advances the state-of-the-art in topic model visualization for document-document and topic-document relations. It brings topic models to life in a way that fosters deep understanding of both corpus and models, allowing users to generate interpretive hypotheses and to suggest further experiments. Such tools are an essential step toward assessing whether topic modeling is a suitable technique for AI and cognitive modeling applications.