Matched samples logistic regression in case-control studies with missing values: when to break the matches

Matched samples logistic regression in case-control studies with missing values: when to break the matches
复制标题

DOI:
10.1177/0962280207082348
复制
发表时间:
2008-12-01
影响因子:
2.3
通讯作者:
Khamis, Harry J.
Khamis, Harry J.
中科院分区:
医学3区
文献类型:
--
作者:
Hansson, Lisbeth;Khamis, Harry J.

文献摘要

被引文献

相似文献

在具有连续协变量的个体病例对照设计中,当存在不同的排除率和不同水平的其他设计参数时,使用模拟数据集来评估条件和无条件最大似然估计。估计过程的有效性是通过方法偏差、估计量的方差、Logistic回归的均方根误差(RMSE)和解释变异的百分比来衡量的。在存在缺失观测的情况下,条件估计的均方根误差高于无条件估计,特别是在1:1匹配的情况下。地层尺寸越小,均方根误差越大,1:1匹配的均方根误差越大。解释变异的百分比似乎对缺失数据不敏感,但条件估计的解释变异百分比通常高于无条件估计。它特别适合1:2匹配的设计。为了最小化RMSE,建议使用较高的匹配率;在这种情况下,条件Logistic回归模型和无条件Logistic回归模型产生的有效性水平相当。为使解释变异百分比最大化,建议采用条件Logistic回归模型的1:2配对设计。
Simulated data sets are used to evaluate conditional and unconditional maximum likelihood estimation in an individual case-control design with continuous covariates when there are different rates of excluded cases and different levels of other design parameters. The effectiveness of the estimation procedures is measured by method bias, variance of the estimators, root mean square error (RMSE) for logistic regression and the percentage of explained variation. Conditional estimation leads to higher RMSE than unconditional estimation in the presence of missing observations, especially for 1:1 matching. The RMSE is higher for the smaller stratum size, especially for the 1:1 matching. The percentage of explained variation appears to be insensitive to missing data, but is generally higher for the conditional estimation than for the unconditional estimation. It is particularly good for the 1:2 matching design. For minimizing RMSE, a high matching ratio is recommended; in this case, conditional and unconditional logistic regression models yield comparable levels of effectiveness. For maximizing the percentage of explained variation, the 1:2 matching design with the conditional logistic regression model is recommended.