Determining sample size for progression criteria for pragmatic pilot RCTs: the hypothesis test strikes back!

Determining sample size for progression criteria for pragmatic pilot RCTs: the hypothesis test strikes back!
复制标题

确定务实试验RCT的进程标准的样本量:假设检验罢工!

DOI:
10.1186/s40814-021-00770-x
复制
发表时间:
2021-02-03
影响因子:
1.7
通讯作者:
Lancaster GA
Lancaster GA
中科院分区:
其他
文献类型:
--
作者:
Lewis M;Bromley K;Sutton CJ;McCray G;Myers HL;Lancaster GA

文献摘要

被引文献

相似文献

目前用于报告初步试验的CONSORT指南不建议对临床结局进行假设检验,因为初步试验检测此类差异的效力不足,而这正是主要试验的目的。它指出,主要评价应侧重于可行性/过程结局的描述性分析(例如招募、依从性、治疗保真度)。虽然不测试临床结果的论点是合理的,但这并不一定适用于可行性/过程结果,其中差异可能很大,并且可以用小样本检测到。此外,试点试验的样本量仍有很多模糊之处。许多试点试验采用“红绿灯”系统,以评估由一套先验标准确定的主试验进展情况。我们构建了一种针对二元可行性结果的假设检验方法,该方法围绕该系统进行,该系统基于对绿色区域(可接受结果)的预期来测试红色区域(不可接受结果),并选择样本量以提供高功效来拒绝红色区域(如果绿色区域为真)。落在红色区域的试验点估计值在统计学上不显著,落在绿色区域的试验点估计值在统计学上显著;琥珀色区域表示可能可接受的结果,统计检验可能显著或不显著。例如,关于治疗保真度,如果我们假设红色区域的上限为50%,绿色区域的下限为75%(分别表示不可接受和可接受的治疗保真度),则在90%把握度和单侧5% α下,分析所需的样本量约为n = 34(仅干预组)。在0-17名参与者(0-50%)范围内观察到的治疗保真度将落入红色区域,并且在统计学上不显著,18-25名(51-74%)落入琥珀色区域,并且可能具有或不具有显著性,26-34名(75-100%)落入绿色区域,并且具有显著性,表明保真度可接受。一般而言,评估几个关键过程结局是否进展至主要试验;复合方法需要评估所有这些结局的进展规则。该方法为试点RCT的过程结局评价提供了一个正式的假设检验和样本量指示框架。在线版本包含补充材料,可通过10.1186/s40814-021-00770-x获得。
The current CONSORT guidelines for reporting pilot trials do not recommend hypothesis testing of clinical outcomes on the basis that a pilot trial is under-powered to detect such differences and this is the aim of the main trial. It states that primary evaluation should focus on descriptive analysis of feasibility/process outcomes (e.g. recruitment, adherence, treatment fidelity). Whilst the argument for not testing clinical outcomes is justifiable, the same does not necessarily apply to feasibility/process outcomes, where differences may be large and detectable with small samples. Moreover, there remains much ambiguity around sample size for pilot trials. Many pilot trials adopt a ‘traffic light’ system for evaluating progression to the main trial determined by a set of criteria set up a priori. We construct a hypothesis testing approach for binary feasibility outcomes focused around this system that tests against being in the RED zone (unacceptable outcome) based on an expectation of being in the GREEN zone (acceptable outcome) and choose the sample size to give high power to reject being in the RED zone if the GREEN zone holds true. Pilot point estimates falling in the RED zone will be statistically non-significant and in the GREEN zone will be significant; the AMBER zone designates potentially acceptable outcome and statistical tests may be significant or non-significant. For example, in relation to treatment fidelity, if we assume the upper boundary of the RED zone is 50% and the lower boundary of the GREEN zone is 75% (designating unacceptable and acceptable treatment fidelity, respectively), the sample size required for analysis given 90% power and one-sided 5% alpha would be around n = 34 (intervention group alone). Observed treatment fidelity in the range of 0–17 participants (0–50%) will fall into the RED zone and be statistically non-significant, 18–25 (51–74%) fall into AMBER and may or may not be significant and 26–34 (75–100%) fall into GREEN and will be significant indicating acceptable fidelity. In general, several key process outcomes are assessed for progression to a main trial; a composite approach would require appraising the rules of progression across all these outcomes. This methodology provides a formal framework for hypothesis testing and sample size indication around process outcome evaluation for pilot RCTs. The online version contains supplementary material available at 10.1186/s40814-021-00770-x.