Power, effects, confidence, and significance: An investigation of statistical practices in nursing research

Power, effects, confidence, and significance: An investigation of statistical practices in nursing research
复制标题

DOI:
10.1016/j.ijnurstu.2013.09.014
复制
发表时间:
2014-05-01
影响因子:
8.1
通讯作者:
Happell, Brenda
Happell, Brenda
中科院分区:
医学1区
文献类型:
--
作者:
Gaskin, Cadeyrn J.;Happell, Brenda

文献摘要

被引文献

相似文献

目的:(a)评估护理研究的统计能力,以检测小、中、大效应量;(b)估计这些研究中实验方面的I型错误率;(c)评估(i)先验功率分析、(ii)效应大小(及其解释)和(iii)可信区间的报告程度。设计:统计回顾。数据来源:根据5年影响因子排名前10位的护理期刊2011年卷发表的论文。综述方法:评估论文的统计功效、实验I型误差的控制、先验功效分析的报告、效应大小的报告和解释,以及置信区间的报告。该分析基于333篇论文,从中确定了10337个推论统计数据。结果:检测小、中、大效应量的中位能力为。40(四分位数间距[IQR] = 0.24 - 0.71), 0.98(IQR = 0.85 -1.00)和1.00 (IQR = 1.00-1.00)。实验中I型错误率的中位数为。54 (iqr = 0.26 - 0.80)。28%的论文报告了先验的功率分析。对Spearman秩相关(100%使用该检验的论文)、泊松回归(100%)、优势比(100%)、Kendall tau相关(100%)、Pearson相关(99%)、逻辑回归(98%)、结构方程建模/验证性因子分析/路径分析(97%)和线性回归(83%)的效应大小进行了常规报告,但双比例z检验的效应大小报告较少(50%)。方差分析/协方差分析/多变量方差分析(18%)、t检验(8%)、Wilcoxon检验(8%)、卡方检验(8%)和Fisher精确检验(7%),未报告符号检验、Friedman检验、McNemar检验、多级模型和Kruskal-Wallis检验。效应量很少被解释。28%的论文报告了置信区间。结论:推论统计在护理研究中的应用、报告和解释有待改进。最重要的是,研究人员应该放弃仅仅根据统计显著性(或不显著性)来解释推断检验结果的误导性做法,而是专注于报告和解释效应大小、置信区间和显著性水平。护理研究人员还需要进行和报告先验能力分析,并解决他们研究中I型实验错误膨胀的问题。爱思唯尔有限公司版权所有版权所有。
Objectives: To (a) assess the statistical power of nursing research to detect small, medium, and large effect sizes; (b) estimate the experiment-wise Type I error rate in these studies; and (c) assess the extent to which (i) a priori power analyses, (ii) effect sizes (and interpretations thereof), and (iii) confidence intervals were reported.Design: Statistical review.Data sources: Papers published in the 2011 volumes of the 10 highest ranked nursing journals, based on their 5-year impact factors.Review methods: Papers were assessed for statistical power, control of experiment-wise Type I error, reporting of a priori power analyses, reporting and interpretation of effect sizes, and reporting of confidence intervals. The analyses were based on 333 papers, from which 10,337 inferential statistics were identified.Results: The median power to detect small, medium, and large effect sizes was .40 (interquartile range [IQR] = .24-.71),.98 (IQR = .85-1.00), and 1.00 (IQR = 1.00-1.00), respectively. The median experiment-wise Type I error rate was .54 (IQR = .26-.80). A priori power analyses were reported in 28% of papers. Effect sizes were routinely reported for Spearman's rank correlations (100% of papers in which this test was used), Poisson regressions (100%), odds ratios (100%), Kendall's tau correlations (100%), Pearson's correlations (99%), logistic regressions (98%), structural equation modelling/confirmatory factor analyses/path analyses (97%), and linear regressions (83%), but were reported less often for two-proportion z tests (50%), analyses of variance/analyses of covariance/multivariate analyses of variance (18%), t tests (8%), Wilcoxon's tests (8%), Chi-squared tests (8%), and Fisher's exact tests (7%), and not reported for sign tests, Friedman's tests, McNemar's tests, multi-level models, and Kruskal-Wallis tests. Effect sizes were infrequently interpreted. Confidence intervals were reported in 28% of papers.Conclusion: The use, reporting, and interpretation of inferential statistics in nursing research need substantial improvement. Most importantly, researchers should abandon the misleading practice of interpreting the results from inferential tests based solely on whether they are statistically significant (or not) and, instead, focus on reporting and interpreting effect sizes, confidence intervals, and significance levels. Nursing researchers also need to conduct and report a priori power analyses, and to address the issue of Type I experiment-wise error inflation in their studies. Crown Copyright (C) 2013 Published by Elsevier Ltd. All rights reserved.