Data Analysis of Asymmetric Structures: Advanced Approaches in Computational Statistics

Data Analysis of Asymmetric Structures: Advanced Approaches in Computational Statistics
复制标题

DOI:
10.1198/tech.2006.s391
复制
发表时间:
2006-05
期刊:
影响因子:
2.5
通讯作者:
A. Chiang
A. Chiang
中科院分区:
工程技术3区
文献类型:
--
作者:
A. Chiang

文献摘要

被引文献

相似文献

这篇文章是适当的一个或两个学期的研究生水平的课程的主题,和先决条件是满意的第一年的研究生统计课程的良好理解。在实现这些程序时,计算机软件/编程当然是必须的,但在这里,作者没有做任何假设,也没有提供具体代码或伪代码的示例(在前言中,推荐使用S-PLUS/R或MATLAB来实现这一目的)。这可以允许文本用于对主题进行更多的理论探索,但如果教师计划使他或她的课程语言特定,则应向学生提供额外的材料。正文以回顾统计推断的基础知识的一章开始,一般来说,后面的章节分为三个主要主题:优化,整合和平滑。每一章后面都有简短但具有挑战性的家庭作业(除了第一章),但附录中没有提供解决方案或提示。在章节中考虑的数据集可以在文本的网站上找到,更新的勘误表列表也可以在那里找到。在大多数情况下,教师可以挑选他或她希望涵盖的章节,而不会失去跳过材料的连续性。第一章,简单地标题为“审查”,奠定了基础与符号约定,并迅速涵盖统计分布,经典/贝叶斯估计,极限理论,和马尔可夫链。第2章,“优化和求解非线性方程”,包含了在单变量和多变量背景下的求根技术的基础知识。所有的标准技术都涵盖了足够的细节,并讨论了收敛的措施,并给予特别关注。在第3章“组合优化”中,事情变得更加复杂,其中引入了NP完全性和NP难问题的概念,以扩大优化应用于非常困难的情况的范围。本主题以第4章结束,第4章介绍了期望最大化(EM)算法及其在基于似然的推理中的作用,沿着了基本方法的各种变体。第二个主题从第5章开始,在那里回顾(希望如此!)的数值积分技术,作者做了很好的工作,在有关的方法,发生在概率计算的问题。这是一个简短的章节,可以很快地涵盖它,但我发现它是第6章题为“模拟和蒙特卡罗积分”的材料的一个很好的序言。本章首先简要描述了蒙特卡罗方法应用于定积分,作者给出了几个应用洞察整个章。当然,任何关于计算统计的课程都应该关注模拟;然而,我应该指出,作者选择放弃设计随机数生成器(RNG)的方法。相反,本章侧重于各种采样方法,并假设读者可以访问RNG的“好软件”。第7章和第8章是关于统计科学最重要的最新进展之一,马尔可夫链蒙特卡罗(MCMC)。在第7章中,我们沿着讨论了Metropolis-Hastings和Gibbs抽样器的收敛性,并对MCMC的更先进的创新进行了概述。第9章关于自助的内容并不属于这三个主题之一,所以我很高兴作者认识到这个重要的工具属于他们的文本。虽然本章只有20页左右的篇幅,但作者们还是抓住了bootstrapping中的大多数重要思想,包括参数和非参数方法、百分位数方法和bootstrapping t。本章以一个关于排列测试的精彩(尽管很短)部分结束,但它为另一个重要的重新排列方法提供了很好的见解(尽管它让我想知道为什么折叠刀也没有包括在本章中)。关于平滑的第三个也是最后一个主题包含关于密度估计、二元平滑技术和多元平滑的三章。在第10章中,介绍了常用的核密度估计,然后是核/带宽的选择和替代方法。第11章讨论了二元平滑,其中有线性平滑,样条和黄土(局部加权散点图平滑)。第12章多元平滑继续预测-响应平滑方法的思想,当响应依赖于预测变量的向量时。如前所述,文本中包含的材料足以教授一年的课程,但对于一个学期的时间框架,作者建议专注于第2章,5-7章和9-11章。如果我教一门关于这个主题的课程,我也会包括第4章的大部分内容。我是否考虑通过这一案文?当然!这是令人难以置信的写得很好,全面;少数狡辩,我与它是可笑的轻微。祝贺作者构建了一个优秀的文本。
This text is appropriate for a one or two semester graduate-level course on the subject, and the prerequisites are satisfied by the good understanding of the first year studies in a graduate statistics curriculum. Computer software/programming is of course a must when implementing the procedures, but here the authors make no assumptions or offer examples with specific code or pseudocode (in the Preface, a recommendation is made for either S–PLUS/R or MATLAB to serve this purpose). This can allow for the text to be used in a more theoretical exploration of the subject matter, but if an instructor plans to make his or her course language-specific, additional materials should be made available to the student. The text begins with a chapter that reviews the basics of statistical inference, and in general the later chapters fall into three major themes: optimization, integration, and smoothing. Short but challenging homework sets follow each chapter (except for the first), but no solutions or hints are offered in an Appendix. The datasets considered in the chapters are available on the text’s website, and an updated errata listing is available there as well. For the most part, an instructor can pick and choose the chapters he or she wishes to cover without any loss of continuity from the skipped material. The first chapter, simply entitled “Review,” lays the foundation with notation conventions and quickly covers statistical distributions, classical/Bayesian estimation, limit theory, and Markov chains. Chapter 2, “Optimization and Solving Nonlinear Equations,” contains the basics of root-finding techniques in both the univariate and multivariate context. All of the standard techniques are covered in sufficient detail, and measures of convergence are discussed and given special attention. Things get significantly more complicated in Chapter 3, “Combinatorial Optimization,” where the notion of NP-completeness and NP-hard problems are introduced to widen the scope of optimization applied to very difficult situations. The theme ends with Chapter 4 on the Expectation Maximization (EM) algorithm and its role in likelihood-based inference, along with assorted variants of the basic method. The second theme begins with Chapter 5 where a review (hopefully!) of numerical integration techniques is presented, and the authors do a nice job in relating the methods to problems that occur in probability calculations. This is a brief chapter and it can be covered very quickly, but I find it to be an excellent preface to the material in Chapter 6 entitled “Simulation and Monte Carlo Integration.” This chapter opens with a brief description of the Monte Carlo method applied to definite integrals, and the authors give insight to several applications throughout the chapter. Certainly, any course on computational statistics should spotlight on simulation; however, I should point out that the authors have chosen to forgo material on the methods used for designing a random number generator (RNG). Instead, the chapter focuses on the various sampling methods and assumes the reader has access to “good software” for an RNG. Chapters 7 and 8 are on one of the most important recent advances in statistical science, Markov chain Monte Carlo (MCMC). In Chapter 7, the Metropolis–Hastings and the Gibbs sampler are covered satisfactorily along with a discussion on convergence, and Chapter 8 surveys the more advanced innovations of MCMC. Chapter 9 on bootstrapping does not fall into one of the three themes, so I’m pleased that the authors recognized that this important tool belongs in their text. Although this chapter is only about 20 pages in length, the authors hit on a majority of the important ideas in bootstrapping, including both parametric and nonparametric methods, percentile approaches, and the bootstrap t. The chapter ends with an excellent, albeit short, section on permutation tests, but it offers good insight to another important resampling method (although it made me wonder why the jackknife was also not included in this chapter). The third and final theme on smoothing contains three chapters on density estimation, techniques for bivariate smoothing, and multivariate smoothing. In Chapter 10, the customary kernel density estimator is introduced, followed with sections on the choice of kernel/bandwidth and alternate methods. Chapter 11 on bivariate smoothing has sections on linear smoothers, splines, and loess (locally weighted scatterplot smoother). Chapter 12 on multivariate smoothing continues the ideas of predictor-response smoothing methods to the case when the response depends on a vector of predictor variables. As mentioned previously, there is enough material contained in the text to teach a year-long course on the topic, but for a one-semester timeframe the authors suggest focusing on Chapters 2, 5–7, and 9–11. If I were teaching a course on this subject, I would also include most of Chapter 4 as well. Would I consider adopting this text? Most definitely! It is incredibly well written and comprehensive; the few quibbles that I have with it are ridiculously minor. Congratulations to the authors for constructing an excellent text.