Evolution of Reporting P Values in the Biomedical Literature, 1990-2015

Evolution of Reporting P Values in the Biomedical Literature, 1990-2015
复制标题

DOI:
10.1001/jama.2016.1952
复制
发表时间:
2016-03-15
影响因子:
120.7
通讯作者:
Ioannidis, John P. A.
Ioannidis, John P. A.
中科院分区:
医学1区
文献类型:
--
作者:
Chavalarias, David;Wallach, Joshua David;Ioannidis, John P. A.

文献摘要

被引文献

相似文献

重要性P值的使用和误用产生了广泛的debats.Objective评估在大规模的P值报告的摘要和全文的生物医学研究文章在过去的25年,并确定如何频繁的统计信息是在其他方式比P值。设计进行自动文本挖掘分析,以提取1990年至2015年期间12 821 790篇MEDLINE摘要和PubMed Central(PMC)中843 884篇摘要和全文文章中报告的P值数据。还评价了151种英文核心临床期刊和PubMed分类的特定文章类型中的P值报告。随机抽取1000篇MEDLINE摘要,手动评估P值和其他类型统计信息的报告;在那些报告经验数据的摘要中,结果文本挖掘在1608736篇MEDLINE摘要中识别出4572043个P值,在1608736篇MEDLINE摘要中识别出3438299个P值,385 393 PMC全文文章。摘要中P值的报告从1990年的7.3%增加到2014年的15.6%。2014年,151种核心临床期刊的33.0%的摘要报告了P值(n = 29 725篇摘要),35.7%的荟萃分析(n = 5620),38.9%的临床试验(n = 4624),54.8%的随机对照试验(n = 13544)和2.4%的综述(n = 71529)。摘要和全文中报告的P值的分布在P值为0.05和0.001或更小时显示出强聚类。随着时间的推移,“最佳”(最具统计学意义)报告的P值适度较小,“最差”(最不具统计学意义)报告的P值变得不那么显著。在具有P值的MEDLINE摘要和PMC全文文章中,96%报告至少1个P值为0.05或更低,PMC全文文章中的比例随时间保持稳定。在1000篇人工审查的摘要中,796篇来自报告经验数据的文章; 15.7%报告了P值(125/796 [95% CI,13.2%-18.4%]),置信区间为2.3%(18/796 [95% CI,1.3%-3.6%]),贝叶斯因子为0%(0/796 [95%CI,0%-0.5%]),效应量为13.9%(111/796 [95% CI,11.6%-16.5%]),其他可能导致P值估计的信息为12.4%(99/796 [95% CI,10.2%-14.9%]),18.1%的患者有关于显著性的定性陈述(181/1000 [95% CI,15.8%-20.6%]);仅1.8%(14/796 [95% CI,1.0%-2.9%])的摘要报告了至少1个效应量和至少1个置信区间。在99篇手动提取的有数据的全文文章中,55篇报道了P值,4篇给出了所有报道的效应量的置信区间,没有一篇使用贝叶斯方法,1篇使用了错误发现率,3篇使用了样本量/把握度计算,5篇指定了主要结局。结论和相关性在对1990-2015年MEDLINE摘要和PMC文章中报道的P值的分析中,随着时间的推移,越来越多的MEDLINE摘要和文章报告了P值,几乎所有具有P值的摘要和文章都报告了统计学显著性结果,并且在亚组分析中,很少有文章包括置信区间、贝叶斯因子或效应量。文章不应报道孤立的P值,而应包括效应量和不确定性度量。
IMPORTANCE The use and misuse of P values has generated extensive debates.OBJECTIVE To evaluate in large scale the P values reported in the abstracts and full text of biomedical research articles over the past 25 years and determine how frequently statistical information is presented in ways other than P values. DESIGN Automated text-mining analysis was performed to extract data on P values reported in 12 821 790 MEDLINE abstracts and in 843 884 abstracts and full-text articles in PubMed Central (PMC) from 1990 to 2015. Reporting of P values in 151 English-language core clinical journals and specific article types as classified by PubMed also was evaluated. A random sample of 1000 MEDLINE abstracts was manually assessed for reporting of P values and other types of statistical information; of those abstracts reporting empirical data, 100 articles were also assessed in full text.MAIN OUTCOMES AND MEASURES P values reported.RESULTS Text mining identified 4 572 043 P values in 1 608 736 MEDLINE abstracts and 3 438 299 P values in 385 393 PMC full-text articles. Reporting of P values in abstracts increased from 7.3% in 1990 to 15.6% in 2014. In 2014, P values were reported in 33.0% of abstracts from the 151 core clinical journals (n = 29 725 abstracts), 35.7% of meta-analyses (n = 5620), 38.9% of clinical trials (n = 4624), 54.8% of randomized controlled trials (n = 13 544), and 2.4% of reviews (n = 71 529). The distribution of reported P values in abstracts and in full text showed strong clustering at P values of .05 and of .001 or smaller. Over time, the "best" (most statistically significant) reported P values were modestly smaller and the "worst" (least statistically significant) reported P values became modestly less significant. Among the MEDLINE abstracts and PMC full-text articles with P values, 96% reported at least 1 P value of .05 or lower, with the proportion remaining steady over time in PMC full-text articles. In 1000 abstracts that were manually reviewed, 796 were from articles reporting empirical data; P values were reported in 15.7%(125/796 [95% CI, 13.2%-18.4%]) of abstracts, confidence intervals in 2.3%(18/796 [95% CI, 1.3%-3.6%]), Bayes factors in 0% (0/796 [95% CI, 0%-0.5%]), effect sizes in 13.9%(111/796 [95% CI, 11.6%-16.5%]), other information that could lead to estimation of P values in 12.4%(99/796 [95% CI, 10.2%-14.9%]), and qualitative statements about significance in 18.1%(181/1000 [95% CI, 15.8%-20.6%]); only 1.8%(14/796 [95% CI, 1.0%-2.9%]) of abstracts reported at least 1 effect size and at least 1 confidence interval. Among 99 manually extracted full-text articles with data, 55 reported P values, 4 presented confidence intervals for all reported effect sizes, none used Bayesian methods, 1 used false-discovery rates, 3 used sample size/power calculations, and 5 specified the primary outcome.CONCLUSIONS AND RELEVANCE In this analysis of P values reported in MEDLINE abstracts and in PMC articles from 1990-2015, more MEDLINE abstracts and articles reported P values over time, almost all abstracts and articles with P values reported statistically significant results, and, in a subgroup analysis, few articles included confidence intervals, Bayes factors, or effect sizes. Rather than reporting isolated P values, articles should include effect sizes and uncertainty metrics.