Toward a diagnostic toolkit for linear models with Gaussian-process distributed random effects.

Toward a diagnostic toolkit for linear models with Gaussian-process distributed random effects.
复制标题

面向具有高斯过程分布随机效应的线性模型的诊断工具包。

DOI:
10.1111/biom.12848
复制
发表时间:
2018
期刊:
影响因子:
1.9
通讯作者:
Banerjee,Sudipto
Banerjee,Sudipto
中科院分区:
数学3区
文献类型:
--
作者:
Bose,Maitreyee;Hodges,JamesS;Banerjee,Sudipto

文献摘要

相似文献

在线性混合模型中,高斯过程(GP)被广泛用作随机效应的分布,其使用限制似然或紧密相关的贝叶斯分析来拟合。这篇文章解决了两个问题。首先,我们提出了理解数据如何确定这些模型中的估计的工具,使用GP的谱基近似,在该近似下,限制似然在形式上与具有身份链接的伽马误差GLM的似然相同。其次,为了检查数据对协变量的支持,并了解添加协变量如何移动GP和拟合误差部分的结果中的变化,我们将线性模型诊断,添加变量图(AVP)应用于原始观测值和数据到谱基函数上的投影。谱域和观测域AVP估计协变量的相同系数,但分别强调低频和高频数据特征,从而分别突出协变量对拟合的GP和误差部分的影响。谱近似适用于在规则网格上观察到的数据;对于在不规则位置观察到的数据,我们建议在应用我们的方法之前将数据平滑到网格。这些方法使用Finley等人的森林生物量数据进行说明。(2008)。
Gaussian processes (GPs) are widely used as distributions of random effects in linear mixed models, which are fit using the restricted likelihood or the closely related Bayesian analysis. This article addresses two problems. First, we propose tools for understanding how data determine estimates in these models, using a spectral basis approximation to the GP under which the restricted likelihood is formally identical to the likelihood for a gamma‐errors GLM with identity link. Second, to examine the data's support for a covariate and to understand how adding that covariate moves variation in the outcomeyout of the GP and error parts of the fit, we apply a linear‐model diagnostic, the added variable plot (AVP), both to the original observations and to projections of the data onto the spectral basis functions. The spectral‐ and observation‐domain AVPs estimate the same coefficient for a covariate but emphasize low‐ and high‐frequency data features respectively and thus highlight the covariate's effect on the GP and error parts of the fit, respectively. The spectral approximation applies to data observed on a regular grid; for data observed at irregular locations, we propose smoothing the data to a grid before applying our methods. The methods are illustrated using the forest‐biomass data of Finley et al. (2008).