Exploration of dimensionality reduction for text visualization

Exploration of dimensionality reduction for text visualization
复制标题

文本可视化降维探索

DOI:
10.1109/cmv.2005.8
复制
发表时间:
2005
期刊:
Coordinated and Multiple Views in Exploratory Visualization (CMV'05)
影响因子:
--
通讯作者:
Elke A. Rundensteiner
Elke A. Rundensteiner
中科院分区:
--
文献类型:
--
作者:
Shiping Huang;M. Ward;Elke A. Rundensteiner

文献摘要

被引文献

相似文献

在文本文档可视化社区中,统计分析工具(例如,主成分分析和多维缩放)和神经计算模型(例如,自组织特征图)已经广泛用于维数降低。通常,结果维度设置为2,因为这有助于绘制结果。这些方法的有效性和有效性在很大程度上取决于所使用的特定数据集和目标应用程序的语义。迄今为止,很少有评价,以评估和比较降维方法和降维过程,无论是数值或经验。本文的重点是提出一种机制,比较和评估降维技术的有效性,在文本文档档案的可视化探索。我们使用多变量可视化技术和交互式可视化探索来研究三个问题:(a)哪种降维技术最好地保留了一组文本文档中的相互关系;(B)结果对输出维度的敏感性如何;(c)我们能否自动从从文档中提取的向量中删除冗余或不重要的词,同时仍然保留大部分信息,从而使降维更有效。为了研究每个问题,我们生成补充维度的基础上,几个降维算法和参数控制这些算法。然后,我们直观地分析和探索的特点,减少维空间内实施的链接,多视图多维视觉探索工具,XmdvTool。我们将导出的尺寸与原始数据中已知的特征进行比较。在确定使用不同数量的产出层面的成果质量时,也采用了定量措施。
In the text document visualization community, statistical analysis tools (e.g., principal component analysis and multidimensional scaling) and neurocomputation models (e.g., self-organizing feature maps) have been widely used for dimensionality reduction. Often the resulting dimensionality is set to two, as this facilitates plotting the results. The validity and effectiveness of these approaches largely depend on the specific data sets used and semantics of the targeted applications. To date, there has been little evaluation to assess and compare dimensionality reduction methods and dimensionality reduction processes, either numerically or empirically. The focus of this paper is to propose a mechanism for comparing and evaluating the effectiveness of dimensionality reduction techniques in the visual exploration of text document archives. We use multivariate visualization techniques and interactive visual exploration to study three problems: (a) Which dimensionality reduction technique best preserves the interrelationships within a set of text documents; (b) What is the sensitivity of the results to the number of output dimensions; (c) Can we automatically remove redundant or unimportant words from the vector extracted from the documents while still preserving the majority of information, and thus make dimensionality reduction more efficient. To study each problem, we generate supplemental dimensions based on several dimensionality reduction algorithms and parameters controlling these algorithms. We then visually analyze and explore the characteristics of the reduced dimensional spaces as implemented within a linked, multiview multidimensional visual exploration tool, XmdvTool. We compare the derived dimensions to features known to be present in the original data. Quantitative measures are also used in identifying the quality of results using different numbers of output dimensions.