A Handbook of Statistical Analyses using S-Plus
A Handbook of Statistical Analyses using S-Plus
复制标题
使用 S-Plus 进行统计分析的手册
DOI:
10.1198/tas.2003.s221
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
J. Manola
中科院分区:
文献类型:
--
作者:
J. Manola
S-Plus is a powerful statistical analysis tool that originatedat Bell Laboratories and has since evolved into a widely used proprietary system. The proprietary software maintains the ef cient command-line mode that is the essence of S, while adding an extensive graphical user interface (GUI). In the second edition of this handbook,Brian Everitt has assembled a collection of statistical analysis problems exemplifying both the command line and the GUI. For those who, like me, nd learning by example to be ef cient and effective, the book is a worthwhile investment. The book consists almost exclusively of worked examples, so it is a good idea for readers to have prior or concurrent exposure to the statistical methods being demonstrated before tackling the examples. The book would be a useful adjunct to a more theoretical textbook in an intermediate or advanced applied-statistics class for quantitative nonstatisticians such as engineers or scientists. It would also be useful as a way for students or applied statisticians to get up to speed quickly on using S-Plus. The book refers to a Web site from which example data les and source code can be downloaded. Once I determined that the example les were actually scripts to create data frames, the download process was smooth and the data loaded without incident. (Some more detail about how to do this or where to locate the downloaded les would have been helpful.) Full text of the datasets and commands is also included in the text. The author makes a case for learning the command-line feature of S-Plus. Since the GUI is capable of invoking only a subset of the functionality available through the command line, learning the command-line language is worthwhile. Each chapter contains an example using the GUI and another using the command line, and concludes with several exercises for the reader to work. Solutions to about half the exercises are provided in the back of the book. Chapters of examples and related concepts progress from simple to complex. Chapter 1 provides an introduction to some essential commands and S-Plus objects; an appendix provides additional material about commands. Chapter 1 then gives an example of an S-Plus session, rst using the GUI and then the command line. Chapter 2 discusses describing data and assessing distributions. It illustrates graphical methods, including scatterplot matrices and normal probability plots as well as descriptive statistical methods. Chapter 3 uses analysis of variance to construct and compare models. The use of Scheff¶e’s method for adjusting for multiple comparisons is illustrated both numerically and graphically using the GUI. Chapter 4 covers multiple regression, with examples illustrating enhancement of scatterplot matrices with lowess and linear predicted ts. The hat matrix is also introduced and computed in Chapter 4. Chapter 5 demonstrates the use of generalized linear models to do logistic regression, with a thoroughlyexplained command-line example followed by a more perfunctory GUI example. Chapter 6 provides a good explanation of some of the issues and options involved in analyzing longitudinal data, followed by the illustration of some of these options using the command-line method. Everitt wisely avoids using the GUI in this chapter. Useful plots are demonstrated. Chapter 7 discusses nonlinear regression and maximum-likelihood estimation, including the use of bootstrapping to obtain con dence intervals. This chapter is also presented entirely in commandline mode. Chapter 8 gives a brief synopsis of survival analysis, followed by examples worked in both command-line and GUI modes. Chapter 9 provides a graphical exploration of a multivariate dataset using the GUI, followed by a command-line demonstration of principal-components analysis. Chapter 10 shows how to do k-means clustering using the GUI, then provides the code for a function to compute the gap statistic. This is the most extensive programming example providedin the book. It is followed by another clustering example using the command-line interface. Chapter 11 concludes with an example illustrating bivariate density estimation using the command line, and discriminant analysis using the GUI.