Using the Whole Cohort in the Analysis of Case-Cohort Data

Using the Whole Cohort in the Analysis of Case-Cohort Data
复制标题

DOI:
10.1093/aje/kwp055
复制
发表时间:
2009-06-01
影响因子:
5
通讯作者:
Kulich, Michal
Kulich, Michal
中科院分区:
医学2区
文献类型:
--
作者:
Breslow, Norman E.;Lumley, Thomas;Kulich, Michal

文献摘要

被引文献

相似文献

病例-队列数据分析通常忽略了未作为病例或对照取样的队列成员的宝贵信息。例如,社区动脉粥样硬化风险(ARIC)研究的调查人员通常只报告其15972名参与者的亚研究样本中10%-15%的受试者数据。剩下的受试者只对分层抽样权重有贡献。在免费的R统计系统(http://cran.r-project.org)中实现的分析方法通过校准或估计来调整采样权值,从而更好地利用数据。通过重新分析来自ARIC冠心病研究的数据和基于国家Wilms肿瘤研究数据的模拟,作者证明这种调整可以显著提高所有受试者已知基线协变量的风险比估计的精度。调整还可以提高部分缺失协变量的精度,这些协变量仅在子研究参与者中已知,当它们的值可以以合理的精度推算其余队列成员时。提供了软件、数据集和教程的链接,详细说明了执行调整后的分析所需的步骤。鼓励流行病学家考虑使用这些方法来提高病例队列分析报告结果的准确性。
Case-cohort data analyses often ignore valuable information on cohort members not sampled as cases or controls. The Atherosclerosis Risk in Communities (ARIC) study investigators, for example, typically report data for just the 10%-15% of subjects sampled for substudies of their cohort of 15,972 participants. Remaining subjects contribute to stratified sampling weights only. Analysis methods implemented in the freely available R statistical system (http://cran.r-project.org) make better use of the data through adjustment of the sampling weights via calibration or estimation. By reanalyzing data from an ARIC study of coronary heart disease and simulations based on data from the National Wilms Tumor Study, the authors demonstrate that such adjustment can dramatically improve the precision of hazard ratios estimated for baseline covariates known for all subjects. Adjustment can also improve precision for partially missing covariates, those known for substudy participants only, when their values may be imputed with reasonable accuracy for the remaining cohort members. Links are provided to software, data sets, and tutorials showing in detail the steps needed to carry out the adjusted analyses. Epidemiologists are encouraged to consider use of these methods to enhance the accuracy of results reported from case-cohort analyses.