An online updating approach for testing the proportional hazards assumption with streams of survival data

An online updating approach for testing the proportional hazards assumption with streams of survival data
复制标题

DOI:
10.1111/biom.13137
复制
发表时间:
2019-11-10
期刊:
影响因子:
1.9
通讯作者:
Schifano, Elizabeth D.
Schifano, Elizabeth D.
中科院分区:
数学3区
文献类型:
--
作者:
Xue, Yishu;Wang, HaiYing;Schifano, Elizabeth D.

文献摘要

被引文献

相似文献

考克斯模型仍然是分析事件发生时间数据的首选,即使是大型数据集也是如此,它依赖于比例风险(PH)假设。当生存数据按块顺序到达时,需要一种快速且最小存储密集型的方法来测试PH假设。我们提出了一种在线更新的方法,更新标准的测试统计量,因为每个新的数据块变得可用,大大减轻了计算负担。在PH的零假设下,所提出的统计量具有与在整个数据流上计算的标准版本相同的渐近分布,其中数据块被合并到一个数据集中。在模拟研究中,当PH假设成立时,基于最新数据块的测试及其变体保持其大小,并且具有很大的能力来检测PH假设的不同违规行为。我们还在模拟中表明,我们的方法可以成功地用于“大数据”,超过了一台计算机的计算资源。该方法是说明与淋巴瘤患者的生存分析,从监测,流行病学和最终结果计划。拟定的检验及时识别了与PH假设的偏差,而基于整个数据的检验未捕获到该偏差。
The Cox model-which remains the first choice for analyzing time-to-event data, even for large data sets-relies on the proportional hazards (PH) assumption. When survival data arrive sequentially in chunks, a fast and minimally storage intensive approach to test the PH assumption is desirable. We propose an online updating approach that updates the standard test statistic as each new block of data becomes available and greatly lightens the computational burden. Under the null hypothesis of PH, the proposed statistic is shown to have the same asymptotic distribution as the standard version computed on an entire data stream with the data blocks pooled into one data set. In simulation studies, the test and its variant based on most recent data blocks maintain their sizes when the PH assumption holds and have substantial power to detect different violations of the PH assumption. We also show in simulation that our approach can be used successfully with "big data" that exceed a single computer's computational resources. The approach is illustrated with the survival analysis of patients with lymphoma cancer from the Surveillance, Epidemiology, and End Results Program. The proposed test promptly identified deviation from the PH assumption, which was not captured by the test based on the entire data.