Unweighted regression models perform better than weighted regression techniques for respondent-driven sampling data: results from a simulation study

Unweighted regression models perform better than weighted regression techniques for respondent-driven sampling data: results from a simulation study
复制标题

DOI:
10.1186/s12874-019-0842-5
复制
发表时间:
2019-10-29
影响因子:
4
通讯作者:
Rotondi, Michael
Rotondi, Michael
中科院分区:
医学3区
文献类型:
--
作者:
Avery, Lisa;Rotondi, Nooshin;Rotondi, Michael

文献摘要

被引文献

相似文献

背景:尚不清楚在分析受访者驱动抽样数据时首选加权回归还是未加权回归。我们的目标是评估各种回归模型的有效性,使用或不使用权重以及各种聚类控制,以根据受访者驱动抽样 (RDS) 收集的数据来估计群体成员资格的风险。方法:使用来自每个群体的 1000 个 RDS 样本,基于连续预测变量的已知分布,对具有不同同质性和患病率水平的 12 个网络群体进行模拟。对每个样本建立加权和未加权二项式和泊松一般线性模型(有或没有各种聚类控制和标准误差调整),并评估有效性、偏差和覆盖率。还估计了人群患病率。结果:在回归分析中,未加权对数链接(泊松)模型在所有人群中保持了名义 I 类错误率。加权二项式回归的偏差很大,I 类错误率高得令人无法接受。使用 RDS 加权逻辑回归估计患病率的覆盖率最高,但在低患病率 (10%) 的情况下建议使用未加权模型。结论:在对 RDS 数据进行回归分析时需要谨慎。即使报告的程度准确,报告的程度低也会过度影响回归估计。因此,建议使用未加权泊松回归。
Background: It is unclear whether weighted or unweighted regression is preferred in the analysis of data derived from respondent driven sampling. Our objective was to evaluate the validity of various regression models, with and without weights and with various controls for clustering in the estimation of the risk of group membership from data collected using respondent-driven sampling (RDS).Methods: Twelve networked populations, with varying levels of homophily and prevalence, based on a known distribution of a continuous predictor were simulated using 1000 RDS samples from each population. Weighted and unweighted binomial and Poisson general linear models, with and without various clustering controls and standard error adjustments were modelled for each sample and evaluated with respect to validity, bias and coverage rate. Population prevalence was also estimated.Results: In the regression analysis, the unweighted log-link (Poisson) models maintained the nominal type-I error rate across all populations. Bias was substantial and type-I error rates unacceptably high for weighted binomial regression. Coverage rates for the estimation of prevalence were highest using RDS-weighted logistic regression, except at low prevalence (10%) where unweighted models are recommended.Conclusions: Caution is warranted when undertaking regression analysis of RDS data. Even when reported degree is accurate, low reported degree can unduly influence regression estimates. Unweighted Poisson regression is therefore recommended.