Evaluating Visual Representations for Topic Understanding and Their Effects on Manually Generated Topic Labels

Evaluating Visual Representations for Topic Understanding and Their Effects on Manually Generated Topic Labels
复制标题

评估主题理解的视觉表示及​​其对手动生成的主题标签的影响

DOI:
--
复制
发表时间:
2017
影响因子:
10.9
通讯作者:
Leah Findlater
Leah Findlater
中科院分区:
人文科学1区
文献类型:
--
作者:
Alison Smith;Tak Yeon Lee;Forough Poursabzi;Jordan L. Boyd;Niklas Elmqvist;Leah Findlater

文献摘要

被引文献

相似文献

概率主题模型是按主题对大型文档集进行索引、摘要和分析的重要工具。然而,促进最终用户对主题的理解仍然是一个开放的研究问题。我们比较标签由用户给出四个主题的可视化技术-单词列表,单词列表与酒吧,词云,网络图-对彼此和自动生成的标签。我们比较的基础是参与者对标签如何描述主题文档的评分。我们的研究分为两个阶段:标记阶段,其中参与者标记可视化的主题;以及验证阶段,其中不同的参与者选择哪个标签最好地描述主题的文档。虽然所有可视化都产生类似的质量标签,但简单的可视化(如单词列表)可以让参与者快速理解主题,而复杂的可视化需要更长的时间,但会暴露出多单词的表达,而简单的可视化会使其模糊不清。自动标签落后于用户创建的标签,但我们的手动标签主题数据集突出了语言模式(例如,上位词、短语),其可用于改进自动主题标记算法。
Probabilistic topic models are important tools for indexing, summarizing, and analyzing large document collections by their themes. However, promoting end-user understanding of topics remains an open research problem. We compare labels generated by users given four topic visualization techniques—word lists, word lists with bars, word clouds, and network graphs—against each other and against automatically generated labels. Our basis of comparison is participant ratings of how well labels describe documents from the topic. Our study has two phases: a labeling phase where participants label visualized topics and a validation phase where different participants select which labels best describe the topics’ documents. Although all visualizations produce similar quality labels, simple visualizations such as word lists allow participants to quickly understand topics, while complex visualizations take longer but expose multi-word expressions that simpler visualizations obscure. Automatic labels lag behind user-created labels, but our dataset of manually labeled topics highlights linguistic patterns (e.g., hypernyms, phrases) that can be used to improve automatic topic labeling algorithms.