Data Analysis of Asymmetric Structures: Advanced Approaches in Computational Statistics
Data Analysis of Asymmetric Structures: Advanced Approaches in Computational Statistics
复制标题
DOI:
10.1198/tech.2006.s391
复制
发表时间:
2006-05
期刊:
影响因子:
2.5
通讯作者:
A. Chiang
中科院分区:
文献类型:
--
作者:
A. Chiang
This text is appropriate for a one or two semester graduate-level course on the subject, and the prerequisites are satisfied by the good understanding of the first year studies in a graduate statistics curriculum. Computer software/programming is of course a must when implementing the procedures, but here the authors make no assumptions or offer examples with specific code or pseudocode (in the Preface, a recommendation is made for either S–PLUS/R or MATLAB to serve this purpose). This can allow for the text to be used in a more theoretical exploration of the subject matter, but if an instructor plans to make his or her course language-specific, additional materials should be made available to the student. The text begins with a chapter that reviews the basics of statistical inference, and in general the later chapters fall into three major themes: optimization, integration, and smoothing. Short but challenging homework sets follow each chapter (except for the first), but no solutions or hints are offered in an Appendix. The datasets considered in the chapters are available on the text’s website, and an updated errata listing is available there as well. For the most part, an instructor can pick and choose the chapters he or she wishes to cover without any loss of continuity from the skipped material. The first chapter, simply entitled “Review,” lays the foundation with notation conventions and quickly covers statistical distributions, classical/Bayesian estimation, limit theory, and Markov chains. Chapter 2, “Optimization and Solving Nonlinear Equations,” contains the basics of root-finding techniques in both the univariate and multivariate context. All of the standard techniques are covered in sufficient detail, and measures of convergence are discussed and given special attention. Things get significantly more complicated in Chapter 3, “Combinatorial Optimization,” where the notion of NP-completeness and NP-hard problems are introduced to widen the scope of optimization applied to very difficult situations. The theme ends with Chapter 4 on the Expectation Maximization (EM) algorithm and its role in likelihood-based inference, along with assorted variants of the basic method. The second theme begins with Chapter 5 where a review (hopefully!) of numerical integration techniques is presented, and the authors do a nice job in relating the methods to problems that occur in probability calculations. This is a brief chapter and it can be covered very quickly, but I find it to be an excellent preface to the material in Chapter 6 entitled “Simulation and Monte Carlo Integration.” This chapter opens with a brief description of the Monte Carlo method applied to definite integrals, and the authors give insight to several applications throughout the chapter. Certainly, any course on computational statistics should spotlight on simulation; however, I should point out that the authors have chosen to forgo material on the methods used for designing a random number generator (RNG). Instead, the chapter focuses on the various sampling methods and assumes the reader has access to “good software” for an RNG. Chapters 7 and 8 are on one of the most important recent advances in statistical science, Markov chain Monte Carlo (MCMC). In Chapter 7, the Metropolis–Hastings and the Gibbs sampler are covered satisfactorily along with a discussion on convergence, and Chapter 8 surveys the more advanced innovations of MCMC. Chapter 9 on bootstrapping does not fall into one of the three themes, so I’m pleased that the authors recognized that this important tool belongs in their text. Although this chapter is only about 20 pages in length, the authors hit on a majority of the important ideas in bootstrapping, including both parametric and nonparametric methods, percentile approaches, and the bootstrap t. The chapter ends with an excellent, albeit short, section on permutation tests, but it offers good insight to another important resampling method (although it made me wonder why the jackknife was also not included in this chapter). The third and final theme on smoothing contains three chapters on density estimation, techniques for bivariate smoothing, and multivariate smoothing. In Chapter 10, the customary kernel density estimator is introduced, followed with sections on the choice of kernel/bandwidth and alternate methods. Chapter 11 on bivariate smoothing has sections on linear smoothers, splines, and loess (locally weighted scatterplot smoother). Chapter 12 on multivariate smoothing continues the ideas of predictor-response smoothing methods to the case when the response depends on a vector of predictor variables. As mentioned previously, there is enough material contained in the text to teach a year-long course on the topic, but for a one-semester timeframe the authors suggest focusing on Chapters 2, 5–7, and 9–11. If I were teaching a course on this subject, I would also include most of Chapter 4 as well. Would I consider adopting this text? Most definitely! It is incredibly well written and comprehensive; the few quibbles that I have with it are ridiculously minor. Congratulations to the authors for constructing an excellent text.