A Bayesian test for excess zeros in a zero-inflated power series distribution

A Bayesian test for excess zeros in a zero-inflated power series distribution
复制标题

零膨胀幂级数分布中多余零点的贝叶斯检验

DOI:
10.1214/193940307000000068
复制
发表时间:
2008
期刊:
arXiv: Statistics Theory
影响因子:
--
通讯作者:
G. Datta
G. Datta
中科院分区:
--
文献类型:
--
作者:
A. Bhattacharya;B. Clarke;G. Datta

文献摘要

被引文献

相似文献

幂级数分布是单参数离散指数族的一个有用的子类,适合于计数数据的建模。零膨胀幂级数分布是幂级数分布与零处简并分布的混合,简并分布的混合概率为$p$。这种分布对于可能有额外零的计数数据的建模很有用。一个问题是混合模型是否可以约简到幂级数部分,对应$p=0$,或者数据中有太多的零,相对于纯幂级数分布的零膨胀是否必须包含在模型中,即$p\geq0$。这个问题之所以困难,部分原因是$p=0$是一个边界点。在认识到参数空间可以扩展以允许$p$为负的基础上,我们提出了这个问题的贝叶斯检验。$p$的负值与$p$作为混合概率的解释不一致,然而,它们表示的分布在物理和概率上都是有意义的。我们将贝叶斯解与两个标准频率测试程序进行比较,发现使用后验概率作为测试统计量在样本量$n$和参数值的最重要范围上比模拟中的得分测试和似然比测试具有稍高的功率。我们的方法在三个真实数据集上也表现良好。
Power series distributions form a useful subclass of one-parameter discrete exponential families suitable for modeling count data. A zero-inflated power series distribution is a mixture of a power series distribution and a degenerate distribution at zero, with a mixing probability $p$ for the degenerate distribution. This distribution is useful for modeling count data that may have extra zeros. One question is whether the mixture model can be reduced to the power series portion, corresponding to $p=0$, or whether there are so many zeros in the data that zero inflation relative to the pure power series distribution must be included in the model i.e., $p\geq0$. The problem is difficult partially because $p=0$ is a boundary point. Here, we present a Bayesian test for this problem based on recognizing that the parameter space can be expanded to allow $p$ to be negative. Negative values of $p$ are inconsistent with the interpretation of $p$ as a mixing probability, however, they index distributions that are physically and probabilistically meaningful. We compare our Bayesian solution to two standard frequentist testing procedures and find that using a posterior probability as a test statistic has slightly higher power on the most important ranges of the sample size $n$ and parameter values than the score test and likelihood ratio test in simulations. Our method also performs well on three real data sets.