Bayesian Data Analysis in Ecology Using Linear Models with R, BUGS, and Stan. F. Korner‐Nievergelt, T. Roth, S. von Felten, J. Guélat, B. Almasi, and P. Korner‐Nievergelt. 2015. Elsevier, London, U.K. 316 pp. $76.39 paperback. ISBN 978‐0‐12‐801370‐0.

Bayesian Data Analysis in Ecology Using Linear Models with R, BUGS, and Stan. F. Korner‐Nievergelt, T. Roth, S. von Felten, J. Guélat, B. Almasi, and P. Korner‐Nievergelt. 2015. Elsevier, London, U.K. 316 pp. $76.39 paperback. ISBN 978‐0‐12‐801370‐0.
复制标题

使用 R、BUGS 和 Stan 的线性模型进行生态学中的贝叶斯数据分析。F. Korner-Nievergelt、T. Roth、S. von Felten、J. Guélat、B. Almasi 和 P. Korner-Nievergelt,2015 年。英国伦敦,316 页,平装本,售价 76.39 美元,ISBN 978-0-12-801370-0。

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
S. Harju
S. Harju
中科院分区:
--
文献类型:
--
作者:
S. Harju

文献摘要

被引文献

相似文献

生态学中的贝叶斯数据分析是一本优秀的统计工具箱书,它提供了使用频率论和贝叶斯方法增加复杂性的生态分析示例。本书面向中级从业者,因为对这两种方法背后的理论和数学的讨论很少。作者还假设读者已经具备 R 程序的应用知识;尽管本书充满了有用的 R 代码片段,但没有介绍 R 的基本工作原理和数据操作步骤。然而,对于那些已经熟悉 R 和一些研究生水平的统计学培训的人(或那些寻求此类培训的人)来说,本书提供了对野生动物和生态数据进行频率论和贝叶斯统计分析的宝贵资源。也许这本书的最大优势在于与野生动物生物学家相关的大量分析,这些分析是在频率论和贝叶斯方法的背景下呈现的。本书首先简要讨论了为什么我们一般需要统计模型(第 1 章),然后是关于 R 中的软件和统计术语和对象的章节(第 2 章:先决条件和词汇)。第三章简要比较和对比了频率论和贝叶斯的哲学和结果。第 4 章(正态线性模型)说明了如何将任一方法应用于最熟悉的模型类型之一:简单线性回归及其变体。第 5 章讨论了似然理论,随后的第 6 章强调评估模型假设。至此,本书深入探讨了与野生动物研究高度相关的日益复杂的模型的章节和示例,包括线性混合效应模型(第7章)、广义线性模型(第8章)、广义线性混合模型(第9章)、模型检查(第10章)、模型选择和多模型推理(第11章)、马尔可夫链蒙特卡罗模拟(第12章)以及空间数据建模等章节。 (第 13 章)。第14章(高级生态模型)快速浏览了野生动物应用中更先进的生态模型,包括栖息地选择的分层模型、零膨胀计数模型、不完善检测的占用估计、无标记个体的领地占用以及标记重新捕获数据的生存估计。其余 3 章讨论评估贝叶斯模型中先验的影响(第 15 章)、进行数据分析的清单(第 16 章)以及如何在分析报告中呈现结果的描述(第 17 章)。全书中反复出现的一个主题是统计分析是一种艺术形式,而不是死记硬背地将工具应用于数据。必须检查假设,必须处理结构问题,最重要的是,必须根据分析的原始目标评估最终统计模型是否合适。对于野生动物生态学中使用的许多类型的模型来说,这既不是一个简单也不直接的任务,鉴于层次模型、混合模型和广义线性模型的复杂性,作者提供了几种评估模型拟合的选项。这本书以强化这一思想并提供读者可以在自己的分析中实施的解决方案而闻名。我对整个“贝叶斯精简版”讨论感到失望。例如,作者确实对频率论和贝叶斯推理进行了非常深刻的比较,尽管过于简短。大多数非统计学家会误解频率论者的 95% 置信区间 (CI) 实际上所描述的内容,认为这意味着真实值有 95% 的可能性位于该范围内。这是不正确的。相反,正如作者所说,“[频率派] CI 的解释是这样的:如果我们在相同条件下使用相同的样本量多次重复研究,那么 95% 的 CI 将包含真实的总体平均值。”换句话说,95% 的时间 95% CI 将包含真实值;我们无法说明在任何特定分析中是否确实如此。相比之下,贝叶斯推理的主要好处之一是我们可以做出直接的概率陈述,例如“我们 95% 确定真实均值在可信区间内”(可信区间是贝叶斯中频率论置信区间的等价物)。除其他外,这种比较很有价值,但过于简短,并且迷失在应用示例的实质内容中。根据本书的意图,比较频率论方法与贝叶斯方法的哲学、实现和推理的更彻底的讨论将被证明是有帮助的。此外,鉴于介绍性方法,读者将不得不寻找其他地方来学习如何利用贝叶斯软件的全部功能和复杂生态模型的推理。尽管如此,本书的简洁性也是其优点。尽管它既不是贝叶斯教科书,也不是频率论教科书,但它提供了通过相当简单的模型应用贝叶斯技术的温和介绍。它还包含许多优秀的建议、警告和解决方案,无论选择何种分析途径,所有从业者都将从中受益。最终,它将在包括我的书架在内的许多书架上占有一席之地,因为它提供了大量的工作示例,并详细参考了相关资源以供进一步探索。
Bayesian Data Analysis in Ecology is an excellent statistical toolbox book that provides examples of ecological analyses that increase in complexity using frequentist and Bayesian methods. The book is intended for intermediate-level practitioners because there is minimal discussion of the theory and mathematics underlying either approach. The authors also assume that the reader already has working knowledge of Program R; although the book is full of useful R code snippets, there is not an introduction to the basic workings and data manipulation steps in R. However, for those already familiar with R and some graduate-level training in statistics (or those searching for such training), this book provides a valuable go-to resource for conducting frequentist and Bayesian statistical analyses of wildlife and ecological data. Perhaps the greatest strength of this book is the sheer number of worked analyses that are relevant to wildlife biologists, presented within the context of frequentist and Bayesian approaches. The book begins with a brief discussion of why we need statistical models in general (Chapter 1), followed by a chapter on software and statistical terms and objects in R (Chapter 2: Prerequisites and Vocabulary). Chapter 3 briefly compares and contrasts the frequentist and Bayesian philosophies and outcomes. Chapter 4 (Normal Linear Models) illustrates how to apply either approach to one of the most familiar types of models: simple linear regression and its variants. Likelihood theory is discussed in Chapter 5, followed by Chapter 6, which emphasizes assessing model assumptions. At this point, the book delves into chapters and worked examples of increasingly complex models that are highly relevant to wildlife studies, including chapters on linear mixed effects models (Chapter 7), generalized linear models (Chapter 8), generalized linear mixed models (Chapter 9), model checking (Chapter 10), model selection and multimodel inference (Chapter 11), Markov Chain Monte Carlo simulation (Chapter 12), and modeling spatial data (Chapter 13). Chapter 14 (Advanced Ecological Models) quickly jumps through more advanced ecological models with wildlife applications, including hierarchical models for habitat selection, zero-inflated count models, occupancy estimation with imperfect detection, territory occupancy with unmarked individuals, and survival estimation from mark-recapture data. The remaining 3 chapters discuss assessing the influence of priors in Bayesian models (Chapter 15), a checklist for conducting data analysis (Chapter 16), and a description of how to present results in analysis write-ups (Chapter 17). One recurring theme throughout the book is that statistical analysis is an art form, not a rote application of tools to data. Assumptions must be checked, structural problems must be dealt with, and above all, final statistical models must be assessed for appropriate fit given the original goals of the analysis. This is neither a simple nor straightforward task for many of the types of models used in wildlife ecology, and the authors provide several options for assessing model fit given the complexity of hierarchical, mixed, and generalized linear models. The book is notable for reinforcing this idea and for providing solutions that the reader can implement in their own analyses. I was disappointed in the “Bayesian-lite” discussion throughout. For example, the authors do make a very poignant, albeit too brief, comparison between frequentist and Bayesian inference. Most non-statisticians misinterpret what the frequentist’s 95% confidence interval (CI) actually is describing, assuming it means that there is a 95% percent chance that the true value lies within this range. This is incorrect. Rather, as the authors state, “The interpretation of the [frequentist] CI is this: If we repeat the study under the same conditions using the same sample size many times, 95% of these CI will contain the true population mean.” In other words, 95% of the time the 95% CI will contain the true value; we can make no statement about whether it does so in any particular analysis. In contrast, one of the major benefits of Bayesian inference is that we can make direct probabilistic statements, such as “we are 95% sure that the true mean is within the Credible Interval” (credible interval is the Bayesian equivalent of the frequentist confidence interval). This comparison, among others, is valuable but was too brief and lost within the meat of the applied examples. A more thorough discussion comparing the philosophy, implementation, and inference of frequentist versus Bayesian approaches would have proven helpful in light of the intent of the book. Additionally, given the introductory approach, readers will have to look elsewhere to learn how to leverage the full power of Bayesian software and inference for complex ecological models. Nevertheless, the book’s brevity is also its strong point. Although it is neither a Bayesian nor frequentist textbook, it provides a gentle introduction to applying Bayesian techniques with fairly simple models. It also contains many excellent suggestions, cautions, and solutions that all practitioners would benefit from regardless of their chosen analytical pathway. And ultimately, it will have a permanent place on many bookshelves, including mine, for the sheer density of worked examples with detailed references to related sources for further exploration.