A Handbook of Statistical Analyses using S-Plus

A Handbook of Statistical Analyses using S-Plus
复制标题

使用 S-Plus 进行统计分析的手册

DOI:
10.1198/tas.2003.s221
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
J. Manola
J. Manola
中科院分区:
--
文献类型:
--
作者:
J. Manola

文献摘要

被引文献

相似文献

S-Plus 是一款功能强大的统计分析工具,起源于贝尔实验室,现已发展成为广泛使用的专有系统。该专有软件保留了 S 本质上的高效命令行模式,同时添加了广泛的图形用户界面 (GUI)。在本手册的第二版中,Brian Everitt 收集了一系列统计分析问题,以命令行和 GUI 为例。对于那些像我一样认为通过实例学习能够高效且有效的人来说,这本书是一项值得的投资。本书几乎全部由工作示例组成,因此读者最好在处理示例之前预先或同时了解所演示的统计方法。对于工程师或科学家等定量非统计学家来说,这本书将是中级或高级应用统计学课程中更具理论性的教科书的有用补充。对于学生或应用统计学家来说,它也有助于快速掌握 S-Plus 的使用。本书引用了一个网站,可以从该网站下载示例数据文件和源代码。一旦我确定示例文件实际上是创建数据帧的脚本,下载过程就很顺利,数据加载也没有发生任何事件。 (有关如何执行此操作或在何处找到下载的文件的更多详细信息会很有帮助。)数据集和命令的全文也包含在文本中。作者给出了学习S-Plus命令行功能的案例。由于 GUI 只能调用命令行可用功能的子集,因此学习命令行语言是值得的。每章都包含一个使用 GUI 的示例和另一个使用命令行的示例,最后提供了一些练习供读者练习。书后提供了大约一半练习的答案。示例和相关概念的章节从简单到复杂。第 1 章介绍了一些基本命令和 S-Plus 对象;附录提供了有关命令的附加材料。然后,第 1 章给出了 S-Plus 会话的示例,首先使用 GUI,然后使用命令行。第 2 章讨论描述数据和评估分布。它说明了图形方法,包括散点图矩阵和正态概率图以及描述性统计方法。第3章使用方差分析来构建和比较模型。使用 GUI 以数字和图形方式说明了如何使用 Scheff¶e 的方法来调整多重比较。第 4 章介绍了多元回归,并举例说明了使用 Lowess 和线性预测 ts 增强散点图矩阵。第 4 章还介绍并计算了帽子矩阵。第 5 章演示了如何使用广义线性模型进行逻辑回归,其中有一个彻底解释的命令行示例,后面是一个更敷衍的 GUI 示例。第 6 章很好地解释了分析纵向数据所涉及的一些问题和选项,然后使用命令行方法说明了其中一些选项。 Everitt 明智地避免在本章中使用 GUI。演示了有用的图。第 7 章讨论非线性回归和最大似然估计,包括使用自举法来获得置信区间。本章也完全​​以命令行模式呈现。第 8 章简要概述了生存分析,随后给出了在命令行和 GUI 模式下运行的示例。第 9 章使用 GUI 对多元数据集进行图形探索,然后进行主成分分析的命令行演示。第 10 章展示了如何使用 GUI 进行 k 均值聚类,然后提供计算间隙统计量的函数的代码。这是本书中提供的最广泛的编程示例。接下来是另一个使用命令行界面的集群示例。第 11 章以一个示例结束,该示例说明了使用命令行进行双变量密度估计以及使用 GUI 进行判别分析。
S-Plus is a powerful statistical analysis tool that originatedat Bell Laboratories and has since evolved into a widely used proprietary system. The proprietary software maintains the efŽ cient command-line mode that is the essence of S, while adding an extensive graphical user interface (GUI). In the second edition of this handbook,Brian Everitt has assembled a collection of statistical analysis problems exemplifying both the command line and the GUI. For those who, like me, Ž nd learning by example to be efŽ cient and effective, the book is a worthwhile investment. The book consists almost exclusively of worked examples, so it is a good idea for readers to have prior or concurrent exposure to the statistical methods being demonstrated before tackling the examples. The book would be a useful adjunct to a more theoretical textbook in an intermediate or advanced applied-statistics class for quantitative nonstatisticians such as engineers or scientists. It would also be useful as a way for students or applied statisticians to get up to speed quickly on using S-Plus. The book refers to a Web site from which example data Ž les and source code can be downloaded. Once I determined that the example Ž les were actually scripts to create data frames, the download process was smooth and the data loaded without incident. (Some more detail about how to do this or where to locate the downloaded Ž les would have been helpful.) Full text of the datasets and commands is also included in the text. The author makes a case for learning the command-line feature of S-Plus. Since the GUI is capable of invoking only a subset of the functionality available through the command line, learning the command-line language is worthwhile. Each chapter contains an example using the GUI and another using the command line, and concludes with several exercises for the reader to work. Solutions to about half the exercises are provided in the back of the book. Chapters of examples and related concepts progress from simple to complex. Chapter 1 provides an introduction to some essential commands and S-Plus objects; an appendix provides additional material about commands. Chapter 1 then gives an example of an S-Plus session, Ž rst using the GUI and then the command line. Chapter 2 discusses describing data and assessing distributions. It illustrates graphical methods, including scatterplot matrices and normal probability plots as well as descriptive statistical methods. Chapter 3 uses analysis of variance to construct and compare models. The use of Scheff¶e’s method for adjusting for multiple comparisons is illustrated both numerically and graphically using the GUI. Chapter 4 covers multiple regression, with examples illustrating enhancement of scatterplot matrices with lowess and linear predicted Ž ts. The hat matrix is also introduced and computed in Chapter 4. Chapter 5 demonstrates the use of generalized linear models to do logistic regression, with a thoroughlyexplained command-line example followed by a more perfunctory GUI example. Chapter 6 provides a good explanation of some of the issues and options involved in analyzing longitudinal data, followed by the illustration of some of these options using the command-line method. Everitt wisely avoids using the GUI in this chapter. Useful plots are demonstrated. Chapter 7 discusses nonlinear regression and maximum-likelihood estimation, including the use of bootstrapping to obtain conŽ dence intervals. This chapter is also presented entirely in commandline mode. Chapter 8 gives a brief synopsis of survival analysis, followed by examples worked in both command-line and GUI modes. Chapter 9 provides a graphical exploration of a multivariate dataset using the GUI, followed by a command-line demonstration of principal-components analysis. Chapter 10 shows how to do k-means clustering using the GUI, then provides the code for a function to compute the gap statistic. This is the most extensive programming example providedin the book. It is followed by another clustering example using the command-line interface. Chapter 11 concludes with an example illustrating bivariate density estimation using the command line, and discriminant analysis using the GUI.