Bayesian Generalized Additive Models for Location, Scale, and Shape for Zero-Inflated and Overdispersed Count Data

Bayesian Generalized Additive Models for Location, Scale, and Shape for Zero-Inflated and Overdispersed Count Data
复制标题

DOI:
10.1080/01621459.2014.912955
复制
发表时间:
2015-03-01
影响因子:
3.7
通讯作者:
Lang, Stefan
Lang, Stefan
中科院分区:
数学1区
文献类型:
--
作者:
Klein, Nadja;Kneib, Thomas;Lang, Stefan

文献摘要

被引文献

相似文献

在应用研究中,阻碍经典泊松对数线性模型分析计数数据的常见问题包括过度离散、与泊松分布相比过多的零点、相关响应以及包含连续协变量、相互作用或空间效应的非线性效应的复杂预报器结构。在位置、规模和形状的广义加性模型的框架内,我们提出了一类适用于零膨胀和过度分散计数数据的贝叶斯广义加性模型,其中可以为计数数据分布的几个参数指定半参数预报器。作为应用功的标准选项,我们考虑零膨胀泊松分布、负二项分布和零膨胀负二项分布。加性预测器规范依赖于不同类型效果的基函数近似,并结合高斯平滑先验。我们发展了基于马尔可夫链蒙特卡罗模拟技术的贝叶斯推理,其中合适的建议密度是基于对完全条件的迭代加权最小二乘逼近来构造的。为了确保推理的实用性,我们考虑了所涉及的理论性质,如关节后验是否适当的问题。在仿真研究中对该方法进行了评估,并将其应用于统计汽车保险中专利引用和索赔频率的数据。对于模型在分布方面的比较,我们认为分位数残差是一种有效的图形工具和评分规则,允许我们量化模型的预测能力。一旦选择了响应分布,就使用偏差信息准则来选择适当的预测器规范。这篇文章的补充材料可以在网上找到。
Frequent problems in applied research preventing the application of the classical Poisson log-linear model for analyzing count data include overdispersion, an excess of zeros compared to the Poisson distribution, correlated responses, as well as complex predictor structures comprising nonlinear effects of continuous covariates, interactions or spatial effects. We propose a general class of Bayesian generalized additive models for zero-inflated and overdispersed count data within the framework of generalized additive models for location, scale, and shape where semiparametric predictors can be specified for several parameters of a count data distribution. As standard options for applied work we consider the zero-inflated Poisson, the negative binomial and the zero-inflated negative binomial distribution. The additive predictor specifications rely on basis function approximations for the different types of effects in combination with Gaussian smoothness priors. We develop Bayesian inference based on Markov chain Monte Carlo simulation techniques where suitable proposal densities are constructed based on iteratively weighted least squares approximations to the full conditionals. To ensure practicability of the inference, we consider theoretical properties like the involved question whether the joint posterior is proper. The proposed approach is evaluated in simulation studies and applied to count data arising from patent citations and claim frequencies in car insurances. For the comparison of models with respect to the distribution, we consider quantile residuals as an effective graphical device and scoring rules that allow us to quantify the predictive ability of the models. The deviance information criterion is used to select appropriate predictor specifications once a response distribution has been chosen. Supplementary materials for this article are available online.