Multi-dimensional penalized hazard model with continuous covariates: applications for studying trends and social inequalities in cancer survival

Multi-dimensional penalized hazard model with continuous covariates: applications for studying trends and social inequalities in cancer survival
复制标题

DOI:
10.1111/rssc.12368
复制
发表时间:
2019-07-22
影响因子:
1.6
通讯作者:
Remontet, Laurent
Remontet, Laurent
中科院分区:
数学3区
文献类型:
--
作者:
Fauvernier, Mathieu;Roche, Laurent;Remontet, Laurent

文献摘要

被引文献

相似文献

描述患者死亡危险的动态变化是癌症流行病学家主要关注的问题。除时间和年龄外,模型中还经常需要包括其他连续协变量。例如,生存趋势分析和社会经济研究分别涉及诊断年份和贫困指数。利用最新的一般平滑模型的理论框架,本文提出了一种对事件间隔时间分析中的风险和超额风险模型进行惩罚的方法。基线风险和协变量的函数形式通过使用惩罚的自然三次回归样条法和相关的二次惩罚来指定。通过形成一个光滑的张量积来处理连续协变量和时变效应之间的相互作用。通过优化拉普拉斯近似边际似然准则或似然交叉验证准则来估计平滑参数。回归参数的估计采用直接最大化生存模型惩罚似然的方法,避免了数据增加和泊松似然方法。通过仿真对所提出的实现方法进行了评估,并应用于实际数据。研究发现,该方法数值稳定、效率高,可用于在总体生存和净生存情境中选择适当的复杂程度,并且简化了模型的建立过程。
Describing the dynamics of patient mortality hazard is a major concern for cancer epidemiologists. In addition to time and age, other continuous covariates have often to be included in the model. For example, survival trend analyses and socio-economic studies deal respectively with the year of diagnosis and a deprivation index. Taking advantage of a recent theoretical framework for general smooth models, the paper proposes a penalized approach to hazard and excess hazard models in time-to-event analyses. The baseline hazard and the functional forms of the covariates were specified by using penalized natural cubic regression splines with associated quadratic penalties. Interactions between continuous covariates and time-dependent effects were dealt with by forming a tensor product smooth. The smoothing parameters were estimated by optimizing either the Laplace approximate marginal likelihood criterion or the likelihood cross-validation criterion. The regression parameters were estimated by direct maximization of the penalized likelihood of the survival model, which avoids data augmentation and the Poisson likelihood approach. The implementation proposed was evaluated on simulations and applied to real data. It was found to be numerically stable, efficient and useful for choosing the appropriate degree of complexity in overall survival and net survival contexts; moreover, it simplified the model building process.