Eliciting priors and relaxing the single causal variant assumption in colocalisation analyses

Eliciting priors and relaxing the single causal variant assumption in colocalisation analyses
复制标题

DOI:
10.1371/journal.pgen.1008720
复制
发表时间:
2020-04-01
期刊:
影响因子:
4.5
通讯作者:
Wallace, Chris
Wallace, Chris
中科院分区:
生物学2区
文献类型:
--
作者:
Wallace, Chris

文献摘要

被引文献

相似文献

确定两个性状是否共享一个遗传原因有助于确定遗传影响疾病或其他性状风险的潜在机制。这样做的一种方法是“coloc”,它在贝叶斯统计框架中更新关于两个性状共享因果变异的机会的先验知识,并观察到遗传关联数据。为了仅使用共同共享的汇总遗传关联数据来做到这一点,该方法做出了某些假设,特别是关于可能构成基因组区域中每个测量特征的遗传因果变异的数量。我们通过几种数据驱动的方法来总结这种技术所需的先验知识,并提出敏感性分析作为检查推理对先验知识的不确定性具有鲁棒性的一种手段。我们还展示了如何在一个地区的因果变异的数量的假设可能会放松,这提高了推理的准确性。水平整合的汇总统计量从不同的GWAS性状可以用来评估证据,为他们共享的遗传因果关系。一种流行的方法是贝叶斯方法,coloc,它只需要GWAS汇总统计量而不需要连锁不平衡估计,现在被常规用于进行数千个性状之间的比较。在这里,我们表明,虽然大多数用户不调整默认的软件值,错误的先验参数可以大大改变后验推理。我们建议数据驱动的方法来获得合理的先验值,并演示如何敏感性分析可以用来评估后验推理的鲁棒性。coloc的灵活性是以不切实际的假设为代价的,即每个性状都有一个单一的因果变异。这种假设可以通过逐步调节来放松,但这需要外部软件和LD矩阵来研究等位基因。我们现在已经在coloc中实现了条件反射,并提出了一种新的替代方法,掩蔽,它不需要LD,并且在因果变量独立时近似条件反射。重要的是,掩蔽可以与条件作用结合使用,其中等位基因对齐的LD估计仅可用于单个性状。我们已经实现了这些发展,在一个新版本的coloc,我们希望这将使更多的知情选择的先验知识,并克服限制的单一因果变量的假设coloc分析。
Author summaryDetermining whether two traits share a genetic cause can be helpful to identify mechanisms underlying genetically-influenced risk of disease or other traits. One method for doing this is "coloc", which updates prior knowledge about the chance of two traits sharing a causal variant with observed genetic association data in a Bayesian statistical framework. To do this using only summary genetic association data that is commonly shared, the method makes certain assumptions, in particular about the number of genetic causal variants that may underlie each measured trait in a genomic region. We walk through several data-driven approaches to summarise the prior knowledge required for this technique, and propose sensitivity analysis as a means of checking that inference is robust to uncertainty about that prior knowledge. We also show how the assumptions about number of causal variants in a region may be relaxed, and that this improves inferential accuracy.Horizontal integration of summary statistics from different GWAS traits can be used to evaluate evidence for their shared genetic causality. One popular method to do this is a Bayesian method, coloc, which is attractive in requiring only GWAS summary statistics and no linkage disequilibrium estimates and is now being used routinely to perform thousands of comparisons between traits. Here we show that while most users do not adjust default software values, misspecification of prior parameters can substantially alter posterior inference. We suggest data driven methods to derive sensible prior values, and demonstrate how sensitivity analysis can be used to assess robustness of posterior inference. The flexibility of coloc comes at the expense of an unrealistic assumption of a single causal variant per trait. This assumption can be relaxed by stepwise conditioning, but this requires external software and an LD matrix aligned to study alleles. We have now implemented conditioning within coloc, and propose a new alternative method, masking, that does not require LD and approximates conditioning when causal variants are independent. Importantly, masking can be used in combination with conditioning where allelically aligned LD estimates are available for only a single trait. We have implemented these developments in a new version of coloc which we hope will enable more informed choice of priors and overcome the restriction of the single causal variant assumptions in coloc analysis.