Approximately unbiased tests of regions using multistep-multiscale bootstrap resampling

Approximately unbiased tests of regions using multistep-multiscale bootstrap resampling
复制标题

DOI:
10.1214/009053604000000823
复制
发表时间:
2004-12-01
影响因子:
4.5
通讯作者:
Shimodaira, H
Shimodaira, H
中科院分区:
数学1区
文献类型:
--
作者:
Shimodaira, H

文献摘要

被引文献

相似文献

基于自举概率的近似无偏检验被认为是具有未知期望参数向量的指数分布族,其中零假设被表示为具有光滑边界的任意形状的区域。这个问题已经讨论了以前在埃夫隆和Tibshirani [安。26(1998)1687-1718],并且通过Efron、Halloran和Holmes的两水平自举法计算具有二阶渐近准确度的校正的p值[Proc.Natl. Acad. Sci. U.S.A. 93(1996)13429-13434]基于Efron的ABC偏差校正[J. Amer Statist. 82(1987)171-185]。我们的论点是他们的渐近理论,其中的几何形状,如有符号的距离和曲率的边界,起着重要的作用。我们给出了另一种计算校正的p值,而不需要在观察值的边界上找到“最近点”,这在两级引导中是必需的,并且在复杂问题中是一个实现负担。其关键思想是改变复制数据集的样本大小与观察数据集的样本大小。对于几个样本量,计算落入该区域的重复样本的频率,然后通过观察频率沿着变化的样本量的变化来计算p值。这是Shimodaira [Systematic Biology 51(2002)492-508]的多尺度自举,对于多变量正态模型,它是三阶精度的。在这里,我们介绍了一个新设计的多步多尺度引导,计算三阶准确的p值的指数分布族。事实上,我们的p值渐近等于Hall的双重引导[The Bootstrap and Edgeworth Expansion(1992)Springer,纽约]和Barndorff-Nielsen [Biometrika 73(1986)307-322]忽略O(n(-3/2))项的修改后的有符号似然比所获得的值,但计算要求较低,不受模型规范的影响。该算法是非常简单的,尽管它背后的理论的复杂性。在简单的例子中说明的p值的差异,并在一个系统的方式显示的自助方法的精度。
Approximately unbiased tests based on bootstrap probabilities are considered for the exponential family of distributions with unknown expectation parameter vector, where the null hypothesis is represented as an arbitrary-shaped region with smooth boundaries. This problem has been discussed previously in Efron and Tibshirani [Ann. Statist. 26 (1998) 1687-1718], and a corrected p-value with second-order asymptotic accuracy is calculated by the two-level bootstrap of Efron, Halloran and Holmes [Proc. Natl. Acad. Sci. U.S.A. 93 (1996) 13429-13434] based on the ABC bias correction of Efron [J. Amer Statist. Assoc. 82 (1987) 171-185]. Our argument is an extension of their asymptotic theory, where the geometry, such as the signed distance and the curvature of the boundary, plays an important role. We give another calculation of the corrected p-value without finding the "nearest point" on the boundary to the observation, which is required in the two-level bootstrap and is an implementational burden in complicated problems. The key idea is to alter the sample size of the replicated dataset from that of the observed dataset. The frequency of the replicates falling in the region is counted for several sample sizes, and then the p-value is calculated by looking at the change in the frequencies along the changing sample sizes. This is the multiscale bootstrap of Shimodaira [Systematic Biology 51 (2002) 492-508], which is third-order accurate for the multivariate normal model. Here we introduce a newly devised multistep-multiscale bootstrap, calculating a third-order accurate p-value for the exponential family of distributions. In fact, our p-value is asymptotically equivalent to those obtained by the double bootstrap of Hall [The Bootstrap and Edgeworth Expansion (1992) Springer, New York] and the modified signed likelihood ratio of Barndorff-Nielsen [Biometrika 73 (1986) 307-322] ignoring O(n(-3/2)) terms, yet the computation is less demanding and free from model specification. The algorithm is remarkably simple despite complexity of the theory behind it. The differences of the p-values are illustrated in simple examples, and the accuracies of the bootstrap methods are shown in a systematic way.