Specifying and implementing nonparametric and semiparametric survival estimators in two-stage (nested) cohort studies with missing case data

Specifying and implementing nonparametric and semiparametric survival estimators in two-stage (nested) cohort studies with missing case data
复制标题

DOI:
10.1198/016214505000000952
复制
发表时间:
2006-06-01
影响因子:
3.7
通讯作者:
Katki, Hormuzd A.
Katki, Hormuzd A.
中科院分区:
数学1区
文献类型:
--
作者:
Mark, Steven D.;Katki, Hormuzd A.

文献摘要

被引文献

相似文献

自1986年以来,我们一直在研究来自中国贲门癌流行率较高地区的一组个体,并进行了许多两阶段研究,以评估各种暴露与这种癌症的相关性。两阶段研究是常用的统计设计。第一阶段涉及观察所有队列成员的结果和可获得的基线协变量信息,第二阶段涉及使用第一阶段的观察结果选择队列的一个子集,用于测量难以获得的暴露。当结果是删失失败时间时,例如在我们的研究中,最常用的设计是病例队列和嵌套病例对照设计。这两种设计的一个局限性是,当某些情况下缺失第二阶段测量时,累积风险的估计值以及生存率和绝对风险的估计值会有偏差。根据我们的经验,这种缺失几乎存在于所有使用生物样本来获得暴露测量值的两阶段研究中(如我们的研究)。在早期的工作中,我们推导出一类非参数和一类半参数累积风险估计的效率和特点是无偏的,无论是否所有的情况下进行测量。在这篇文章中,我们将这两个类的数学推导的介绍限制在研究设计和分析的重要方面。我们分析了一项关于幽门螺杆菌感染与贲门癌发病之间关系的两阶段研究的数据。我们讨论了为什么我们故意只对25%的癌症病例进行抽样的实质性原因。通过模拟,我们证明了在精度上的实质性变化存在于每个类内的无偏估计之间,并表达了这些差异的起源,研究人员熟悉的参数。我们描述了如何预先存在的知识,这些参数可以用来提高估计精度,并详细说明了具体的策略,构建这样的估计。实现这些估计器的R语言计算机代码可向作者索取。
Since 1986, we have been studying a cohort of individuals from a region in China with epidemic rates of gastric cardia cancer and have conducted numerous two-stage studies to assess the association of various exposures with this cancer. Two-stage studies are a commonly used statistical design. Stage one involves observing the outcomes and accessible baseline covariate information on all cohort members, and stage two involves using the stage one observations to select a subset of the cohort for measurements of exposures that are difficult to obtain. When the outcomes are censored failure times, such as in our studies, the most common designs used are the case-cohort and nested case-control designs. One limitation of both these designs is that the estimators of the cumulative hazards, and hence survivals and absolute risks, are biased when some cases are missing the stage two measurements. In our experience, such missingness is present in virtually all two-stage studies that (like ours) use biological specimens to obtain exposure measurements. In earlier work we derived and characterized the efficiency of a class of nonparametric and a class of semiparametric cumulative hazard estimators that are unbiased regardless of whether or not all cases are measured. In this article we limit the presentation of the mathematical derivation of these two classes to aspects important to study design and analysis. We analyze data from a two-stage study that we conducted on the association of Helicobacter pylori infection with incident gastric cardia cancers. We discuss the substantive reasons why we deliberately sampled only 25% of the available cancer cases. Through simulations, we demonstrate that substantial variation in precision exists between unbiased estimators within each class, and express the origin of these differences in terms of parameters familiar to investigators. We describe how preexistent knowledge about these parameters can be used to increase estimator precision, and detail specific strategies for constructing such estimators. Computer code in R that implements these estimators is available from the authors on request.