A Two-stage Linear Mixed Model (TS-LMM) for Summary-data-based Multivariable Mendelian Randomization.

A Two-stage Linear Mixed Model (TS-LMM) for Summary-data-based Multivariable Mendelian Randomization.
复制标题

用于基于汇总数据的多变量孟德尔随机化的两阶段线性混合模型 (TS-LMM)。

DOI:
10.1101/2023.04.25.23289099
复制
发表时间:
2023
期刊:
medRxiv : the preprint server for health sciences
影响因子:
--
通讯作者:
Ding,Ming
Ding,Ming
中科院分区:
--
文献类型:
--
作者:
Ding,Ming

文献摘要

相似文献

多变量孟德尔随机化(MVMR)方法提供了一种应用全基因组汇总统计来评估多个风险因素对疾病结局的同时因果影响的策略。与假设无遗传多效性(遗传变异仅与一个风险因素相关)的单变量MR方法相比,MVMR允许遗传变异与多个风险因素相关,并通过将风险因素作为多个变量的汇总统计量纳入回归模型来模拟多效性。在这里,我们提出了一个两阶段的线性混合模型(TS-LMM)的MVMR,占方差的汇总统计不仅在结果中,而且在所有的危险因素。在第一阶段,我们应用线性混合模型将疾病汇总统计量中的方差视为固定/随机效应,同时考虑由于连锁不平衡(LD)引起的遗传变异之间的协方差。特别地,我们使用迭代重加权最小二乘算法来获得随机效应的估计。在第二阶段,我们通过应用测量误差校正方法同时考虑多个风险因素的汇总统计量的方差,该方法考虑了遗传变异之间的LD和风险因素汇总统计量之间的相关性。我们比较了我们的MVMR方法与其他方法在模拟研究。当大多数工具变量(IV)都很强时,我们的模型对真实因果关系的覆盖率最高,检测显著因果关系的能力最高,并且在一系列不同相关性的场景中识别零因果效应的假阳性率最低(弱,强)之间的风险因素和LD的汇总统计量的遗传变异(弱LD [γ2≤0.1],中度LD [0.1< γ2≤0.5])。当强IV的比例减少时,我们的模型显示出与MVMR-Egger和MVMR-IVW相当的性能。在风险因素之间存在相关性的情况下,我们的模型的更准确的推断支持潜在的广泛应用于组学数据,这些数据通常是多维的和相关的,如在应用于寿命的决定因素中所示,其中我们的方法从一组10个脂蛋白胆固醇测量中提名了一个特定的显著脂蛋白亚组分用于因果关联。我们的模型对相关性结构的鲁棒性表明,在实践中,我们可以在IV的选择中允许中等LD,从而可能以更有效的方式利用全基因组汇总数据。我们的模型是在R语言的'TS_LMM'宏中实现的。
Multivariable Mendelian randomization (MVMR) methods provide a strategy for applying genome-wide summary statistics to assess simultaneous causal effects of multiple risk factors on a disease outcome. In contrast to univariate MR methods that assumes no horizonal pleiotropy (genetic variants only associate with one risk factor), MVMR allows for genetic variants associate with multiple risk factors and models pleiotropy by including summary statistics with risk factors as multiple variables into the regression model. Here, we propose a two-stage linear mixed model (TS-LMM) for MVMR that accounts for variance of summary statistics not only in outcome, but also in all of the risk factors. In stage I, we apply linear mixed model to treat variance in summary statistics of disease as fixed-/random-effects, while accounting for covariance between genetic variants due to linkage disequilibrium (LD). Particularly, we use an iteratively re-weighted least squares algorithm to obtain estimates for the random-effects. In stage II, we account for variance in summary statistics of multiple risk factors simultaneously by applying measurement error correction methods that take into consideration LD between genetic variants and correlation between summary statistics of risk factors. We compared our MVMR approach to other approaches in a simulation study. When most of the instrumental variables (IVs) were strong, our model generated the highest coverage of true causal associations, the highest power of detecting significant causal associations, and the lowest false positive rate of identifying null causal effect for a range of scenarios that varied correlation (weak, strong) between summary statistics of risk factors and LD among genetic variants (weak LD [γ2≤0.1], moderate LD [0.1< γ2≤0.5]). When the proportion of strong IVs was reduced, our model showed performances comparable to MVMR-Egger and MVMR-IVW. The more accurate inference of our model in the presence of correlation among risk factors supports potential wide application to -omics data that are commonly multi dimensional and correlated, as shown in application to determinants of longevity, where our method nominated a specific significant lipoprotein subfraction for causal association from a panel of 10 lipoprotein cholesterol measures. The robustness of our model to correlation structure suggests that in practice we can allow moderate LD in selection of IVs, thereby potentially leveraging genome-wide summary data in a more effective manner. Our model is implemented in ‘TS_LMM’ macro in R.