Comparing effect sizes across variables: generalization without the need for Bonferroni correction

Comparing effect sizes across variables: generalization without the need for Bonferroni correction
复制标题

DOI:
10.1093/beheco/ark005
复制
发表时间:
2006-07-01
期刊:
影响因子:
2.4
通讯作者:
Garamszegi, Laszlo Zsolt
Garamszegi, Laszlo Zsolt
中科院分区:
环境科学与生态学2区
文献类型:
--
作者:
Garamszegi, Laszlo Zsolt

文献摘要

被引文献

相似文献

行为生态学的研究经常调查几个特征,然后应用多种统计测试来发现它们的成对关联。传统上,这种方法需要调整个体的显著性水平,因为随着进行更多的统计检验,I型错误发生的可能性就越大(即,当H 0为真时拒绝H 0)(Rice 1989)。Bonferroni校正,降低临界P值为每个特定的测试的基础上进行的测试的数量是经常使用,以减少与多重比较相关的问题(Cabin和米切尔2000年)。然而,这个过程大大增加了犯II类错误的风险,因为它导致当H 0为假时不拒绝H 0的高风险。为了达到80%的统计功效,有必要拥有巨大的样本量来检测中等(r ½ 0.3或d ½ 0.5;意义Cohen 1988)或小(r ½ 0.1或d ½ 0.2;意义Cohen 1988)强度效应(例如,对于双样本t检验,分别说N ½ 128或N ½ 788),但在研究行为时,样本量往往有限。因此,在生态学和行为生态学领域中严格应用邦弗罗尼校正的做法受到了数学和逻辑上的批评(Wright 1992; Benjamini and Hochberg 1995; Perneger 1998; Moran 2003; Nakagawa 2004)。作为一种潜在的解决方案,Wright(1992)和钱德勒(1995)主张通过选择高于通常接受的5%的实验误差率来避免功率的牺牲损失,这导致不同类型的误差之间的平衡。作为另一种选择,研究人员可能更感兴趣的是控制错误拒绝的零假设的比例,所谓的错误发现率,而不是控制家族错误率(Benjamini和Hochberg,1995)。虽然这种方法可以在大量重复试验中提高功效,但很少应用于生态学研究(Garcia 2003,2004)。最近,Nakagawa(2004)建议报告所有潜在关系的效应量和置信区间(CI),以便读者判断结果的生物学重要性并减少发表偏倚。由于检验的功效较低,预计大多数研究的关系都不显著,这被认为是难以发表的。这种困难通常被认为是导致行为生态学家选择性地报告数据(Moran 2003; Nakagawa 2004)。从科学和伦理的角度来看,从出版物中遗漏不显著的结果是不可取的,这使得Bonferroni调整存在问题。值得注意的是,比较已发表和未发表研究的代表性样本的效应量的直接检验显示,生物学文献中没有发表偏倚的证据(Koricheva 2003; Møller et al. 2005)。然而,在这方面,
Studies in behavioral ecology often investigate several traits and then apply multiple statistical tests to discover their pairwise associations. Traditionally, such approaches require the adjustment of individual significance levels because as more statistical tests are performed the greater the likelihood that Type I errors are committed (ie, rejecting H0 when it is true)(Rice 1989). Bonferroni correction that lowers the critical P values for each particular test based on the number of tests to be performed is frequently used to reduce problems associated with multiple comparisons (Cabin and Mitchell 2000). However, this procedure dramatically increases the risk of committing Type II errors as it results in a high risk of not rejecting a H0 when it is false. To reach 80% statistical power, it is necessary to have huge sample sizes to detect medium (r ¼ 0.3 or d ¼ 0.5; sensu Cohen 1988) or small (r ¼ 0.1 or d ¼ 0.2; sensu Cohen 1988) strength effects (eg, say N ¼ 128 or N ¼ 788, respectively, for a 2-sample t-test), but sample size is often limited when studying behavior. The strict application of Bonferroni correction in the field of ecology and behavioral ecology has therefore been criticized for mathematical and logical reasons (Wright 1992; Benjamini and Hochberg 1995; Perneger 1998; Moran 2003; Nakagawa 2004). As a potential solution, Wright (1992) and Chandler (1995) advocated that the sacrificial loss of power can be avoided by choosing an experimentwise error rate higher than the usually accepted 5%, which results in a balance between different types of errors. As another alternative, the researcher might be more interested in controlling the proportion of erroneously rejected null hypotheses, the socalled false discovery rate, than in controlling for familywise error rate (Benjamini and Hochberg, 1995). Although this approach allows for increased power in large series of repeated tests, it is rarely applied in ecological studies (Garcia 2003, 2004).Recently, Nakagawa (2004) suggested reporting effect sizes together with confidence intervals (CIs) for all potential relationships to allow the readers to judge the biological importance of the results and to reduce publication bias. Due to the low power of the tests, the majority of investigated relationships are expected to be nonsignificant, which is thought to make publication difficult. Such difficulty is generally assumed to cause behavioral ecologists to selectively report data (Moran 2003; Nakagawa 2004). The omission of nonsignificant results from publications is undesirable for both scientific and ethical reasons, which makes Bonferroni adjustment problematic. It is noteworthy that direct tests comparing effect sizes of representative samples of published and unpublished studies showed no evidence of publication bias in the biological literature (Koricheva 2003; Møller et al. 2005). However,