Dismantling the Fragility Index: A demonstration of statistical reasoning

Dismantling the Fragility Index: A demonstration of statistical reasoning
复制标题

DOI:
10.1002/sim.8689
复制
发表时间:
2020-08-11
影响因子:
2
通讯作者:
Potter, Gail E.
Potter, Gail E.
中科院分区:
医学3区
文献类型:
--
作者:
Potter, Gail E.

文献摘要

被引文献

相似文献

脆弱性指数被引入作为P值的补充,以总结试验结果证据的统计强度。脆弱性指数(FI)定义为两个治疗组规模相等的试验,具有二分或至事件发生时间结局,计算为治疗组中将Fisher精确检验的P值移动到0.05阈值以上所需的从无事件转换为事件的最小数量。由于该指数缺乏明确的概率动机,其解释对消费者来说具有挑战性。我们澄清了什么FI可能捕获分别考虑两种情况:(a)什么FI捕获数学概率模型是正确的时,(B)如何以及FI捕获违反概率模型假设。通过计算后验概率的治疗效果,我们表明,当概率模型是正确的,FI不适当地惩罚小试验使用更少的事件比大型试验,以达到相同的显着性水平。分析表明,对于在没有偏见的情况下进行的实验,FI会促进对概率的不正确直觉,这在其他地方没有注意到,必须消除。我们说明了FI的能力,量化偏离模型假设和上下文的FI概念在当前的辩论中的零假设显著性检验范式的缺点。总而言之,FI造成的混乱比它解决的要多,并且没有促进统计思维。我们建议不要使用它。相反,建议进行敏感性分析,以量化和传达试验结果的稳健性。
The Fragility Index has been introduced as a complement to theP-value to summarize the statistical strength of evidence for a trial's result. The Fragility Index (FI) is defined in trials with two equal treatment group sizes, with a dichotomous or time-to-event outcome, and is calculated as the minimum number of conversions from nonevent to event in the treatment group needed to shift theP-value from Fisher's exact test over the .05 threshold. As the index lacks a well-defined probability motivation, its interpretation is challenging for consumers. We clarify what the FI may be capturing by separately considering two scenarios: (a) what the FI is capturing mathematically when the probability model is correct and (b) how well the FI captures violations of probability model assumptions. By calculating the posterior probability of a treatment effect, we show that when the probability model is correct, the FI inappropriately penalizes small trials for using fewer events than larger trials to achieve the same significance level. The analysis shows that for experiments conducted without bias, the FI promotes an incorrect intuition of probability, which has not been noted elsewhere and must be dispelled. We illustrate shortcomings of the FI's ability to quantify departures from model assumptions and contextualize the FI concept within current debate around the null hypothesis significance testing paradigm. Altogether, the FI creates more confusion than it resolves and does not promote statistical thinking. We recommend against its use. Instead, sensitivity analyses are recommended to quantify and communicate robustness of trial results.