Generalized Linear Models

Generalized Linear Models
复制标题

DOI:
10.1198/004017002320256422
复制
发表时间:
2002-08
期刊:
影响因子:
2.5
通讯作者:
E. Ziegel
E. Ziegel
中科院分区:
工程技术3区
文献类型:
--
作者:
E. Ziegel

文献摘要

被引文献

相似文献

这是第一本由与生物科学无关的作者撰写的关于广义线性模型的书。副标题为“在工程和科学中的应用”,这本书的作者都主要专注于工程统计。第一个作者最近出版了几个版本的Walpole, Myers, and Myers(1998),最后一个由Ziegel(1999)报道。第二位作者有几个版本的Montgomery和Runger(1999),最近由Ziegel(2002)报道。所有作者都是著名的建模专家。前两位作者合作编写了一本应用建模的开创性著作(Myers and Montgomery 2002), Ziegel(2002)最近对其进行了修订。最后两位作者合作出版了一本关于回归分析的书的最新版本(Montgomery, Peck, and Vining, 2001),由Gray(2002)报道,第一个作者已经出版了他自己的回归分析书的多个版本(Myers 1990),其中最新的版本由Ziegel(1991)报道。Conklin(2002)报道的Hosmer和Lemeshow(2000)是一本具有类似目标和更具体地关注逻辑回归的可比书籍,它假定回归分析的背景,并从广义线性模型开始。这里的序言(第11页)表明了同样的要求,但仍然以100页关于线性和非线性回归的材料开始。其中大部分内容可能是给这本书的读者的评论。第2章,“线性回归模型”,以50页熟悉的关于多元回归的估计、推断和诊断检查的材料开始。该方法是非常传统的,包括使用正式的假设检验。在工业环境中,使用p值作为风险加权决策的一部分通常更合适。教学方法包括公式和计算演示,尽管最后用Minitab说明计算。关于最大似然估计,缩放残差和加权最小二乘的不太熟悉的材料为广义线性模型的后续估计方法提供了更具体的背景。这篇评论并不是要贬低别人。作者在本章中为任何实践者提供了大量有用的金块。它读起来非常有趣。第3章,“非线性回归模型”,可以说不是一个回顾,因为回归分析课程经常对非线性模型给予短暂的关注。本章以一个关于非线性模型线性化用于参数估计的陷阱的很好的例子开始。它继续有效地平衡关于计算理论基础的明确陈述与应用和演示它们的使用。再次给出了极大似然估计的细节,并讨论了加权回归估计和广义回归估计。第4章的题目是“逻辑回归模型和泊松回归模型”。逻辑回归为广义线性模型提供了基本模型。先前的加权回归的发展被用来激发逻辑模型中参数的最大似然估计。提供了代数细节。在线性模型的开发中,一些细节被放在附录中。除了在若干场合与上述有关回归的材料相联系之外,作者还将他们的发展与下一章有关整个广义线性模型家族的内容联系起来。他们讨论分数函数,方差-协方差矩阵,沃尔德推理,似然推理,偏差和过度分散。对标准计算机软件中提供的值进行了仔细的解释,这里是SAS中的PROC LOGISTIC。当在过度分散和非同质方差,或偏差分析和方差分析之间进行类比时,清晰地认识到这本书从熟悉的回归概念开始的价值。作者依靠泊松回归方法与逻辑回归方法的相似性,并大多给出了泊松回归的插图。这些在SAS中使用PROC GENMOD。这本书没有给出任何产生结果的SAS代码。其中两个例子说明了设计的实验和建模。其中包括对子集选择和过度色散调整的讨论。在第5章“广义线性模型族”中提高了演示的数学水平。首先,作者在指数分布下统一了前两章。关于广义线性模型(GLMs)的形式结构、似然方程、拟似然、伽玛分布族和作为链接的幂函数的材料是本书中一些最先进的材料。大部分计算细节都归入附录。对残差的讨论使人回到更实际的角度,关于伽马分布应用程序的两个长示例为如何将这些材料付诸实践提供了很好的指导。一个例子是与使用线性回归与响应的对数变换的对比,另一个是与前一章中使用不同的链接函数的比较。第6章考虑了纵向和类似研究的广义估计方程(GEEs)。本章的前半部分介绍了该方法,后半部分通过五个不同的例子演示了它的应用。一般情况的基础是首先用响应呈正态分布的情况和恒等联系建立起来的。说明了相关结构的重要性,给出了迭代估计过程,讨论了系数尺度参数和标准误差的估计。然后将这些过程推广到指数族分布和拟似然估计。其中两个例子是来自生物统计学应用的标准重复测量插图,但最后三个插图都是工业应用的有趣的重新工作。PROC GENMOD中的GEE计算用于解释与受试者的多个测量或随机化限制有关的相关性。算例表明,考虑相关结构可以得出不同的结论。第7章“GLM的进一步进展和应用”讨论了几个额外的主题。这些是glm的实验设计,渐近结果,筛选实验的分析,数据转换,过程均值和方差的建模,以及广义加性模型。关于实验设计的材料更多是话语性的,而不是说明性的,因此也有些理论性。类似的评论也适用于关于渐近结果的质量的讨论,这在各种模拟研究的报告中有点过于沉湎。关于筛选和数据转换实验的例子再次是对熟悉的工业例子的分析的重新工作,这是作者对使用GLM工具包所产生的热情的另一个明显动机。人们可以希望,后续版本将类似地包含新的例子,这将使作者在本章中扩展关于广义加性模型和其他主题的材料。指定自己去评论一本我知道自己会喜欢读的书,这是作为编辑的回报之一。我读了麦卡拉格和奈尔德(1989)的两个版本,这是由舒恩梅耶(1992)评论过的。那本书读起来没意思。Myers、Montgomery和Vining的明显热情,以及他们对许多例子的依赖,作为他们教学的主要焦点,使得广义线性模型读起来很有趣。在任何应用科学领域工作的每一位统计学家都应该购买这本书,并体验用这些新方法处理熟悉活动所带来的兴奋。
This is the Ž rst book on generalized linear models written by authors not mostly associated with the biological sciences. Subtitled “With Applications in Engineering and the Sciences,” this book’s authors all specialize primarily in engineering statistics. The Ž rst author has produced several recent editions of Walpole, Myers, and Myers (1998), the last reported by Ziegel (1999). The second author has had several editions of Montgomery and Runger (1999), recently reported by Ziegel (2002). All of the authors are renowned experts in modeling. The Ž rst two authors collaborated on a seminal volume in applied modeling (Myers and Montgomery 2002), which had its recent revised edition reported by Ziegel (2002). The last two authors collaborated on the most recent edition of a book on regression analysis (Montgomery, Peck, and Vining (2001), reported by Gray (2002), and the Ž rst author has had multiple editions of his own regression analysis book (Myers 1990), the latest of which was reported by Ziegel (1991). A comparable book with similar objectives and a more speciŽ c focus on logistic regression, Hosmer and Lemeshow (2000), reported by Conklin (2002), presumed a background in regression analysis and began with generalized linear models. The Preface here (p. xi) indicates an identical requirement but nonetheless begins with 100 pages of material on linear and nonlinear regression. Most of this will probably be a review for the readers of the book. Chapter 2, “Linear Regression Model,” begins with 50 pages of familiar material on estimation, inference, and diagnostic checking for multiple regression. The approach is very traditional, including the use of formal hypothesis tests. In industrial settings, use of p values as part of a risk-weighted decision is generally more appropriate. The pedagologic approach includes formulas and demonstrations for computations, although computing by Minitab is eventually illustrated. Less-familiar material on maximum likelihood estimation, scaled residuals, and weighted least squares provides more speciŽ c background for subsequent estimation methods for generalized linear models. This review is not meant to be disparaging. The authors have packed a wealth of useful nuggets for any practitioner in this chapter. It is thoroughly enjoyable to read. Chapter 3, “Nonlinear Regression Models,” is arguably less of a review, because regression analysis courses often give short shrift to nonlinear models. The chapter begins with a great example on the pitfalls of linearizing a nonlinear model for parameter estimation. It continues with the effective balancing of explicit statements concerning the theoretical basis for computations versus the application and demonstration of their use. The details of maximum likelihood estimation are again provided, and weighted and generalized regression estimation are discussed. Chapter 4 is titled “Logistic and Poisson Regression Models.” Logistic regression provides the basic model for generalized linear models. The prior development for weighted regression is used to motivate maximum likelihood estimation for the parameters in the logistic model. The algebraic details are provided. As in the development for linear models, some of the details are pushed into an appendix. In addition to connecting to the foregoing material on regression on several occasions, the authors link their development forward to their following chapter on the entire family of generalized linear models. They discuss score functions, the variance-covariance matrix, Wald inference, likelihood inference, deviance, and overdispersion. Careful explanations are given for the values provided in standard computer software, here PROC LOGISTIC in SAS. The value in having the book begin with familiar regression concepts is clearly realized when the analogies are drawn between overdispersion and nonhomogenous variance, or analysis of deviance and analysis of variance. The authors rely on the similarity of Poisson regression methods to logistic regression methods and mostly present illustrations for Poisson regression. These use PROC GENMOD in SAS. The book does not give any of the SAS code that produces the results. Two of the examples illustrate designed experiments and modeling. They include discussion of subset selection and adjustment for overdispersion. The mathematic level of the presentation is elevated in Chapter 5, “The Family of Generalized Linear Models.” First, the authors unify the two preceding chapters under the exponential distribution. The material on the formal structure for generalized linear models (GLMs), likelihood equations, quasilikelihood, the gamma distribution family, and power functions as links is some of the most advanced material in the book. Most of the computational details are relegated to appendixes. A discussion of residuals returns one to a more practical perspective, and two long examples on gamma distribution applications provide excellent guidance on how to put this material into practice. One example is a contrast to the use of linear regression with a log transformation of the response, and the other is a comparison to the use of a different link function in the previous chapter. Chapter 6 considers generalized estimating equations (GEEs) for longitudinal and analogous studies. The Ž rst half of the chapter presents the methodology, and the second half demonstrates its application through Ž ve different examples. The basis for the general situation is Ž rst established using the case with a normal distribution for the response and an identity link. The importance of the correlation structure is explained, the iterative estimation procedure is shown, and estimation for the scale parameters and the standard errors of the coefŽ cients is discussed. The procedures are then generalized for the exponential family of distributions and quasi-likelihood estimation. Two of the examples are standard repeated-measures illustrations from biostatistical applications, but the last three illustrations are all interesting reworkings of industrial applications. The GEE computations in PROC GENMOD are applied to account for correlations that occur with multiple measurements on the subjects or restrictions to randomizations. The examples show that accounting for correlation structure can result in different conclusions. Chapter 7, “Further Advances and Applications in GLM,” discusses several additional topics. These are experimental designs for GLMs, asymptotic results, analysis of screening experiments, data transformation, modeling for both a process mean and variance, and generalized additive models. The material on experimental designs is more discursive than prescriptive and as a result is also somewhat theoretical. Similar comments apply for the discussion on the quality of the asymptotic results, which wallows a little too much in reports on various simulation studies. The examples on screening and data transformations experiments are again reworkings of analyses of familiar industrial examples and another obvious motivation for the enthusiasm that the authors have developed for using the GLM toolkit. One can hope that subsequent editions will similarly contain new examples that will have caused the authors to expand the material on generalized additive models and other topics in this chapter. Designating myself to review a book that I know I will love to read is one of the rewards of being editor. I read both of the editions of McCullagh and Nelder (1989), which was reviewed by Schuenemeyer (1992). That book was not fun to read. The obvious enthusiasm of Myers, Montgomery, and Vining and their reliance on their many examples as a major focus of their pedagogy make Generalized Linear Models a joy to read. Every statistician working in any area of applied science should buy it and experience the excitement of these new approaches to familiar activities.