Understanding text corpora with multiple facets

Understanding text corpora with multiple facets
复制标题

DOI:
10.1109/vast.2010.5652931
复制
发表时间:
2010-12
期刊:
2010 IEEE Symposium on Visual Analytics Science and Technology
影响因子:
--
通讯作者:
Lei Shi;Furu Wei;Shixia Liu;Li Tan;Xiaoxiao Lian;Michelle X. Zhou
Lei Shi;Furu Wei;Shixia Liu;Li Tan;Xiaoxiao Lian;Michelle X. Zhou
中科院分区:
其他
文献类型:
--
作者:
Lei Shi;Furu Wei;Shixia Liu;Li Tan;Xiaoxiao Lian;Michelle X. Zhou

文献摘要

被引文献

相似文献

文本可视化成为一个越来越重要的研究课题,因为对许多人和企业来说,理解大规模文本信息的需求被证明是必不可少的。然而,由于文本的非结构化和高维性,设计有效的视觉隐喻来表示大型文本语料库仍然是非常具有挑战性的。在本文中,我们提出了一个数据模型,可以用来表示大多数的文本语料库。这样的数据模型包含四种基本类型的方面:时间、类别、内容(非结构化)和结构化方面。为了理解这样的数据模型的语料库,我们开发了一个混合可视化相结合的趋势图与标签云。我们用四个独立的视觉维度对四种类型的数据面进行编码。为了帮助人们发现进化和相关模式,我们还开发了几种视觉交互方法,允许人们通过一个或多个方面交互地分析文本。最后,我们提出了两个案例研究,以证明我们的解决方案的有效性,支持多方面的文本语料库的视觉分析。
Text visualization becomes an increasingly more important research topic as the need to understand massive-scale textual information is proven to be imperative for many people and businesses. However, it is still very challenging to design effective visual metaphors to represent large corpora of text due to the unstructured and high-dimensional nature of text. In this paper, we propose a data model that can be used to represent most of the text corpora. Such a data model contains four basic types of facets: time, category, content (unstructured), and structured facet. To understand the corpus with such a data model, we develop a hybrid visualization by combining the trend graph with tag-clouds. We encode the four types of data facets with four separate visual dimensions. To help people discover evolutionary and correlation patterns, we also develop several visual interaction methods that allow people to interactively analyze text by one or more facets. Finally, we present two case studies to demonstrate the effectiveness of our solution in support of multi-faceted visual analysis of text corpora.