A new method for choosing sample size for confidence interval-based inferences

A new method for choosing sample size for confidence interval-based inferences
复制标题

DOI:
10.1111/1541-0420.00068
复制
发表时间:
2003-09-01
期刊:
影响因子:
1.9
通讯作者:
Stewart, PW
Stewart, PW
中科院分区:
数学3区
文献类型:
--
作者:
Jiroutek, MR;Muller, KE;Stewart, PW

文献摘要

被引文献

相似文献

科学家经常需要检验假设并构建相应的置信区间。在设计一项研究来检验特定的零假设时,传统方法会导致样本量足够大以提供足够的统计功效。相比之下,基于构建置信区间的传统方法导致样本大小可能控制区间的宽度。无论采用哪种方法,样本量过大都会浪费资源或引入伦理问题,都是不可取的。这项工作的动机是担心现有的样本量方法常常使科学家难以实现他们的实际目标。我们关注涉及代表真实自然状态的固定、未知标量参数的情况。置信区间的宽度定义为(随机)上限和下限之间的差。如果观察到的置信区间宽度小于先验选择的固定常数,则称事件宽度发生。如果感兴趣的参数包含在观察到的置信区间上限和下限之间,则称发生事件有效性。如果置信区间排除参数的空值,则称发生事件拒绝。在我们看来,科学家们常常隐含地寻求这三者同时发生:宽度、有效性和拒绝。新的结果表明,忽略拒绝或宽度(以及有效性较低)通常会提供所有三个事件同时发生的概率较低的样本量。我们建议在选择确定样本量的标准时同时考虑所有三个事件。我们为具有高斯误差和固定预测变量的一般线性模型中的任何标量(平均)参数提供了新的理论结果。包括方便的计算形式,以及说明我们的方法的数值示例。
Scientists often need to test hypotheses and construct corresponding confidence intervals. In designing a study to test a particular null hypothesis, traditional methods lead to a sample size large enough to provide sufficient statistical power. In contrast, traditional methods based on constructing a confidence interval lead to a sample size likely to control the width of the interval. With either approach, a sample size so large as to waste resources or introduce ethical concerns is undesirable. This work was motivated by the concern that existing sample size methods often make it difficult for scientists to achieve their actual goals. We focus on situations which involve a fixed, unknown scalar parameter representing the true state of nature. The width of the confidence interval is defined as the difference between the (random) upper and lower bounds. An event width is said to occur if the observed confidence interval width is less than a fixed constant chosen a priori. An event validity is said to occur if the parameter of interest is contained between the observed upper and lower confidence interval bounds. An event rejection is said to occur if the confidence interval excludes the null value of the parameter. In our opinion, scientists often implicitly seek to have all three occur: width, validity, and rejection. New results illustrate that neglecting rejection or width (and less so validity) often provides a sample size with a low probability of the simultaneous occurrence of all three events. We recommend considering all three events simultaneously when choosing a criterion for determining a sample size. We provide new theoretical results for any scalar (mean) parameter in a general linear model with Gaussian errors and fixed predictors. Convenient computational forms are included, as well as numerical examples to illustrate our methods.