Binomial Confidence Intervals and Contingency Tests: Mathematical Fundamentals and the Evaluation of Alternative Methods

Binomial Confidence Intervals and Contingency Tests: Mathematical Fundamentals and the Evaluation of Alternative Methods
复制标题

DOI:
10.1080/09296174.2013.799918
复制
发表时间:
2013-08-01
影响因子:
1.4
通讯作者:
Wallis, Sean
Wallis, Sean
中科院分区:
人文科学4区
文献类型:
--
作者:
Wallis, Sean

文献摘要

被引文献

相似文献

许多统计方法依赖于基于简单近似值的概率的基本数学模型,该模型同时是知名的,但经常被误解。对二项式分布的正常近似是一系列统计检验和方法的基础,包括计算准确的置信区间,拟合和应急测试的良好性,线条和模型拟合以及基于这些测试的计算方法。一个普遍的错误是假设,由于可能正态分布的人口中真实值的误差的可能分布,因此对于观察的错误,可以说相同。该论文分为两个部分:基本原理和评估。首先,我们使用三种初始方法检查置信区间的估计:WALD(正常)间隔,Wilson分数间隔和确切的Clopper-Pearson二项式间隔。尽管可以直接从公式中计算前两个,但必须通过计算搜索近似二项式间隔,并且计算上的昂贵。但是,此间隔提供了最精确的显着性测试,因此将构成我们以后评估的基线。我们还考虑了另外两种改进:在间隔(也需要搜索)中采用对数可能性,以及添加连续性校正的效果。第二,我们在三个测试范式中评估了每种方法。这些是单比例的间隔,或2 x 1的拟合测试优点,在共同的2 x 2偶然性测试中进行了两个变化。我们通过从业者策略评估每种方法的性能。由于标准建议是在预期近似失败的情况下回到确切的二项式测试中,因此我们报告了一个实例的比例,其中一个测试在不相同的精确测试时获得了重要的结果,反之亦然,反之亦然。可能的值。我们证明了最佳方法是基于Wilson Interval或Yates测试的连续性校正版本,并且对测试弱点的通常信念具有误导性。通常提出的对数可能会改进,令人失望地表现出色。最后,我们注意到,在这个精确级别上,我们可以根据自变量分区数据是否将两种类型的2 2测试区分开,并为其使用而做出实用的建议。
Many statistical methods rely on an underlying mathematical model of probability based on a simple approximation, one that is simultaneously well-known and yet frequently misunderstood. The Normal approximation to the Binomial distribution underpins a range of statistical tests and methods, including the calculation of accurate confidence intervals, performing goodness of fit and contingency tests, line- and model-fitting, and computational methods based upon these. A common mistake is in assuming that, since the probable distribution of error about the true value in the population is approximately Normally distributed, the same can be said for the error about an observation.This paper is divided into two parts: fundamentals and evaluation. First, we examine the estimation of confidence intervals using three initial approaches: the Wald (Normal) interval, the Wilson score interval and the exact Clopper-Pearson Binomial interval. Whereas the first two can be calculated directly from formulae, the Binomial interval must be approximated towards by computational search, and is computationally expensive. However this interval provides the most precise significance test, and therefore will form the baseline for our later evaluations. We also consider two further refinements: employing log-likelihood in intervals (also requiring search) and the effect of adding a continuity correction.Second, we evaluate each approach in three test paradigms. These are the single proportion interval or 2 x 1 goodness of fit test, and two variations on the common 2 x 2 contingency test. We evaluate the performance of each approach by a practitioner strategy. Since standard advice is to fall back to exact Binomial tests in conditions when approximations are expected to fail, we report the proportion of instances where one test obtains a significant result when the equivalent exact test does not, and vice versa, across an exhaustive set of possible values.We demonstrate that optimal methods are based on continuity-corrected versions of the Wilson interval or Yates' test, and that commonly-held beliefs about weaknesses of tests are misleading. Log-likelihood, often proposed as an improvement on , performs disappointingly. Finally we note that at this level of precision we may distinguish two types of 2 2 test according to whether the independent variable partitions data into independent populations, and we make practical recommendations for their use.