SIGNIFICANCE TESTING IN THE COMPARISON OF SURVIVAL CURVES FROM CLINICAL-TRIALS OF CANCER-TREATMENT

SIGNIFICANCE TESTING IN THE COMPARISON OF SURVIVAL CURVES FROM CLINICAL-TRIALS OF CANCER-TREATMENT
复制标题

DOI:
10.1016/0277-5379(86)90133-1
复制
发表时间:
1986-11-01
期刊:
EUROPEAN JOURNAL OF CANCER & CLINICAL ONCOLOGY
影响因子:
--
通讯作者:
HAYBITTLE, JL
HAYBITTLE, JL
中科院分区:
其他
文献类型:
--
作者:
HAYBITTLE, JL

文献摘要

被引文献

相似文献

对数秩检验(Peto and Peto 1972)现在被广泛用于比较需要长时间随访的癌症治疗随机临床试验的生存数据。当一组的死亡率始终以一定比例超过另一组时,即所谓的比例危险情况,测试是最优的。有时使用的替代检验是Gehan对Wilcoxon秩和检验的概括(Gehan 1965)以及Peto和Peto(1972)和Prentice(1978)随后对其进行的修改。其中,后者更适用于审查数据(Prentice and Marek 1979),并且,正如Lee et al.(1975)在比较威布尔分布建模的生存曲线的模拟实验中所示,它在非比例风险情况下可能表现更好。类似地,Fleming等人(1980)在比较早期随访时间差异最大的生存曲线时,证明了logrank与Wilcoxon试验相比的有效性损失,Harrington和Fleming(1982)也表明,当风险比在时间为零时达到最大值,并随着随访时间的增加而平稳地趋于一致时,也存在类似的损失。两种测试性能之间存在差异的原因是,Wilcoxon统计量的计算根据事件发生时的估计存活率对观察到的事件和预期事件之间的差异进行加权,而log-rank计算在所有事件时间给出相同的权重(Tarone和Ware 1977)。因此,Wilcoxon检验更重视在随访早期出现的差异。比例风险模型是否最适合许多癌症治疗试验,这可能会受到质疑,因为对照组可能经常包含一个长期幸存者或“治愈”患者的子集,而正在测试的新疗法不太可能对这个子集的生存率有任何改善。例如,Nissen-Meyer(1979)假设辅助治疗的效果
The log-rank test (Peto and Peto 1972) is now widely used for comparing survival data from randomised clinical trials of cancer treatment that require prolonged follow-up. The test is optimal when the death rate in one group consistently exceeds that in the other group by a given proportion, the so-called proportional hazards situation. Alternative tests that are sometimes used are Gehan's generalisation of the Wilcoxon rank sum test (Gehan 1965) and its subsequent modifications by Peto and Peto (1972) and by Prentice (1978). Of these, the latter is to be preferred with censored data (Prentice and Marek 1979), and, as shown by Lee et al.(1975) in a simulation experiment comparing survival curves modelled on Weibull distributions, it may perform better in a nonproportional hazards situation. Similarly, Fleming et al.(1980) have demonstrated the loss of power of the logrank compared with that of the Wilcoxon test in comparing survival curves where the greatest differences occur at early follow-up times, and Harrington and Fleming (1982) have shown a similar loss when the hazard ratio is a maximum at time zero and decreases smoothly towards unity as follow-up increases. The reason for the difference between the performance of the two tests is that the calculation of the Wilcoxon statistic weights the differences between observed and expected events according to the estimated survival at the time of the event, whereas the log-rank calculation gives equal weights at all event times (Tarone and Ware 1977). Thus, the Wilcoxon test gives more weight to differences which appear early in follow-up.It may be questioned whether the proportional-hazards model is the most appropriate one for many trials of cancer therapy, since the control arm may often contain a subset of long-term survivors or'cured'patients, and the new therapy being tested is unlikely to effect any improvement of survival in this subset. For example, Nissen-Meyer (1979) has postulated that the effect of adjuvant therapy