On the use of zero-inflated and Hurdle models for modeling vaccine adverse event count data

On the use of zero-inflated and Hurdle models for modeling vaccine adverse event count data
复制标题

DOI:
10.1080/10543400600719384
复制
发表时间:
2006-07-01
影响因子:
1.1
通讯作者:
Plikaytis, B. D.
Plikaytis, B. D.
中科院分区:
医学4区
文献类型:
--
作者:
Rose, C. E.;Martin, S. W.;Plikaytis, B. D.

文献摘要

被引文献

相似文献

我们比较了疫苗不良事件计数数据的几种建模策略,其中数据具有过零性和异方差特征。计数数据通常使用泊松和负二项(NB)回归进行建模,但零膨胀和栅栏模型在这种情况下可能更有利。这里我们比较了泊松、负二项(NB)、零膨胀泊松(ZIP)、零膨胀负二项(ZINB)、泊松栅栏(PH)和负二项栅栏(NBH)模型的拟合。一般而言,对于公共卫生研究,我们可以将零膨胀模型概念化为允许风险人群和非风险人群出现零。相比之下,障碍模型可能被概念化为只有高危人群的零。我们的结果表明,对于我们的数据,ZINB和NBH模型是首选的,但这些模型在拟合方面是不可区分的。假设Poisson和NB模型因过多的零而不充分,在零膨胀和栅栏模型框架之间进行选择,通常应基于研究的设计和目的。如果研究的目的是推理,那么就应该考虑建模框架。例如,如果研究设计导致对具有结构零和样本零的终点进行计数,则通常零膨胀建模框架更合适,而相反,如果感兴趣的终点通过设计仅展示样本零(例如,处于风险中的参与者),则栏模型框架通常是优选的。相反,如果这项研究的主要目的是开发一个预测模型,那么零膨胀和跨栏建模框架都应该足够。
We compared several modeling strategies for vaccine adverse event count data in which the data are characterized by excess zeroes and heteroskedasticity. Count data are routinely modeled using Poisson and Negative Binomial ( NB) regression but zero-inflated and hurdle models may be advantageous in this setting. Here we compared the fit of the Poisson, Negative Binomial ( NB), zero-inflated Poisson ( ZIP), zero-inflated Negative Binomial (ZINB), Poisson Hurdle (PH), and Negative Binomial Hurdle (NBH) models. In general, for public health studies, we may conceptualize zero-inflated models as allowing zeroes to arise from at-risk and not-at-risk populations. In contrast, hurdle models may be conceptualized as having zeroes only from an at-risk population. Our results illustrate, for our data, that the ZINB and NBH models are preferred but these models are indistinguishable with respect to fit. Choosing between the zero-inflated and hurdle modeling framework, assuming Poisson and NB models are inadequate because of excess zeroes, should generally be based on the study design and purpose. If the study's purpose is inference then modeling framework should be considered. For example, if the study design leads to count endpoints with both structural and sample zeroes then generally the zero-inflated modeling framework is more appropriate, while in contrast, if the endpoint of interest, by design, only exhibits sample zeroes ( e. g., at risk participants) then the hurdle model framework is generally preferred. Conversely, if the study's primary purpose it is to develop a prediction model then both the zero-inflated and hurdle modeling frameworks should be adequate.