A Modern Theory of Factorial Design

A Modern Theory of Factorial Design
复制标题

DOI:
10.1198/tech.2007.s517
复制
发表时间:
2007-08
期刊:
影响因子:
2.5
通讯作者:
Jason L. Loeppky
Jason L. Loeppky
中科院分区:
工程技术3区
文献类型:
--
作者:
Jason L. Loeppky

文献摘要

被引文献

相似文献

统计图表的设计和实施阶段在成功应对大数据集带来的挑战方面发挥着非常重要的作用。第4章集中讨论交互式图形在大型数据集上的应用,以展示它们如何通过使用一些复杂的技术(如查询、缩放、链接、缩放和排序)来帮助提高图形在提供更多信息方面的能力和潜力。本章最后介绍了在处理大型且复杂的数据集时,应如何改进现有数据分析工具的交互能力。关于应用程序的特定图形类型的第5-11章构成了这本名为“应用程序”的书的第二部分。第五章考察了涉及大量案例和大量类别的多变量分类数据的马赛克图及其变化。本章还展示了使用重新排序和红线标记等技术为数据中的内容生成更清晰的图像是多么有价值。第6章集中于有用的可视化技术,称为旋转曲线图,可用于探索多变量连续数据,并概述了为此目的开发和使用的软件。第7章很好地介绍了平行坐标图的平滑修改版本,该版本使用颜色画笔来分隔类别,并可以轻松地处理三维以上的许多观测。本章包含了一些数学知识来解释新的修改后的平行图背后的逻辑,并通过一个例子展示了一些复杂的操作,如轴的重新排序和重新缩放,在获得大型数据集的有用表示时可以发挥的作用。第8章通过集中于在合理的时间限制内绘制最佳布局和产生信息性显示之间的差异来检查网络。由于关注时间和质量之间的差异,本章的重点是探索和发现,让读者通过使用各种布局算法来揭示隐藏在非常大的图表中的不寻常的特征。第9章考察了已知对探索小数据集中的关键模式有用的树。本章将树背后的逻辑概括到大型数据集,重点是组合和总结从复杂数据集可以生成的大量树中获得的信息。本章详细讨论了一些创新的可视化方法,如波动图和树状图,以从大数据集中获取更多信息。第10章通过将“老鼠和大象”图形应用于互联网分组数据,集中讨论了与其相关的一些具有挑战性的可视化问题。本章提供的方法,如有偏采样和分位数窗口,用于处理大型和复杂数据集的可视化,揭示了互联网流量数据集中一些有趣的结构。最后一章,也就是第11章,使用真实的数据集将本书前面几章中描述的可视化方法付诸实施,以说明可视化在数据分析过程中的价值。这是一本有价值的书,对于所有需要实用指导的研究人员来说,他们需要借助各种可视化方法来探索他们的大型数据集。本书讨论了几个应用程序,并对这些方法进行了易于阅读的解释,使它们很容易适用于任何人。我们推荐这本书给每个有兴趣通过各种最先进的可视化技术发现隐藏在大型数据集中的结构的人。由于作者在整本书中都使用了“一百万”这个词作为大型数据集的“有用的象征性目标”,我们也认为阅读这本书的第三页将是有用的,以了解一个多世纪前弗朗西斯·高尔顿对一百万可能是什么样子的想法。
that the design and implementation stages of statistical graphics play a very important role in dealing successfully with the challenges arising from large datasets. Chapter 4 concentrates on the application of interactive graphics to large datasets to show how they can help improve the power and potential of graphics in supplying more information by employing some sophisticated techniques such as querying, scaling, linking, zooming, and sorting. This chapter ends with some remarks about what should be done to improve interaction capabilities of existing data analysis tools when dealing with large and complex datasets. Chapters 5–11 that are on particular types of graphics for applications constitute the second part of the book titled as “Applications.” Chapter 5 examines the mosaic plots and their variations for multivariate categorical data involving both large numbers of cases and large numbers of categories. This chapter also demonstrates how valuable it is to use techniques like reordering and redmarking to produce clearer images of what is in the data. Chapter 6 concentrates on useful visualization techniques, known as rotating plots, that can be used to explore multivariate continuous data, and overviews the software developed and used for this purpose. Chapter 7 provides a very nice introduction to a smooth modified version of the parallel coordinate plot that uses color brushing to separate categories and can easily handle many observations in more than three dimensions. This chapter contains some mathematics to explain the logic behind the new modified parallel plot and shows by an example the role some sophisticated actions such as reordering and rescaling of axes can play in obtaining useful representations of large datasets. Chapter 8 examines the networks by concentrating on the difference between drawing optimal layouts and producing informative displays within a reasonable time limit. Because of focusing on the difference between time and quality, the emphasis in this chapter is on exploration and discovery that let the reader to reveal unusual features hidden within very large graphs by using a variety of layout algorithms. Chapter 9 examines trees that are known to be useful for exploring crucial patterns in the small datasets. This chapter generalizes the logic behind the trees to the large datasets by focusing on the task of combining and summarizing the information obtained from large numbers of trees complex datasets can produce. The behavior of some innovative visualization methods, such as fluctuation diagrams and treemaps, to get more information from large datasets is considered in detail in this chapter. Chapter 10 concentrates on some challenging visualization problems related to the “Mice and Elephants” graphic by applying it to the internet packet data. The methods offered in this chapter, such as biased sampling and quantile windows, to deal with the visualizations of large and complex datasets are shown to uncover some interesting structures in the internet traffic dataset. The last chapter, Chapter 11, puts the visualization methods described in the previous chapters of the book into action using a real dataset to illustrate how valuable visualization can be in the process of data analysis. This is a valuable book for all the researchers who need practical guidance to explore their large datasets by the help of a variety of visualization methods. Several applications discussed throughout the book and easy-to-read explanations of the methods make them easy to apply for anyone. We recommend this book for everyone interested in discovering structures hidden in the large datasets by a variety of state-of-the-art visualization techniques. Since the authors use the word “a million” throughout the book as a “useful symbolic target” for large datasets, we also think that it is going to be useful to read the third page of the book to learn what Francis Galton thought more than a century ago about what a million might look like.