Exploring High-D Spaces with Multiform Matrices and Small Multiples.

Exploring High-D Spaces with Multiform Matrices and Small Multiples.
复制标题

使用多元矩阵和小倍数探索高维空间。

DOI:
10.1109/infvis.2003.1249006
复制
发表时间:
2003
期刊:
IEEE Conference on Information Visualization : an International Conference on Computer Visualization & Graphics, proceedings ... IEEE Conference on Information Visualization
影响因子:
--
通讯作者:
Lengerich,Gene
Lengerich,Gene
中科院分区:
--
文献类型:
--
作者:
Maceachren,Alan;Dai,Xiping;Hardisty,Frank;Guo,Diansheng;Lengerich,Gene

文献摘要

相似文献

我们介绍了一种集成了信息可视化、探索性数据分析(EDA)和地理可视化的多种方法的多变量数据可视化分析方法。该方法利用在Geovista Studio中实现的基于组件的体系结构来构建灵活、多视图、紧密(但一般地)协调的EDA工具包。这个工具包以三种基本方式构建在小倍数和散点图矩阵背后的传统思想之上。首先,我们发展了一个一般的,多元的,二元的矩阵和一个互补的,多元的,二元的小多重图,其中不同的二元表示形式可以组合使用。我们用矩阵和小倍数演示了这种方法的灵活性,这些矩阵和小倍数通过以下组合来描述多变量数据:散点图、双变量图和空间填充显示。其次,我们应用条件熵的度量来(A)从高维数据集中识别可能显示感兴趣的关系的变量,以及(B)在矩阵或小的多个显示中生成这些变量的默认顺序。第三,我们添加了条件,这是一种动态查询/过滤,其中使用补充(未显示的)变量将视图限制在显示的变量上。条件反射允许从分析中去除一个或多个已被充分理解的变量的影响,使剩余变量之间的关系更容易探索。我们通过对癌症诊断和死亡率数据及其相关协变量和风险因素的分析,说明了这种方法所实现的个体和组合功能。
We introduce an approach to visual analysis of multivariate data that integrates several methods from information visualization, exploratory data analysis (EDA), and geovisualization. The approach leverages the component-based architecture implemented in GeoVISTA Studio to construct a flexible, multiview, tightly (but generically) coordinated, EDA toolkit. This toolkit builds upon traditional ideas behind both small multiples and scatterplot matrices in three fundamental ways. First, we develop a general, multiform, bivariate matrix and a complementary multiform, bivariate small multiple plot in which different bivariate representation forms can be used in combination. We demonstrate the flexibility of this approach with matrices and small multiples that depict multivariate data through combinations of: scatterplots, bivariate maps, and space-filling displays. Second, we apply a measure of conditional entropy to (a) identify variables from a high-dimensional data set that are likely to display interesting relationships and (b) generate a default order of these variables in the matrix or small multiple display. Third, we add conditioning, a kind of dynamic query/filtering in which supplementary (undisplayed) variables are used to constrain the view onto variables that are displayed. Conditioning allows the effects of one or more well understood variables to be removed form the analysis, making relationships among remaining variables easier to explore. We illustrate the individual and combined functionality enabled by this approach through application to analysis of cancer diagnosis and mortality data and their associated covariates and risk factors.