Hierarchical modeling for estimating relative risks of rare genetic variants: properties of the pseudo-likelihood method.

Hierarchical modeling for estimating relative risks of rare genetic variants: properties of the pseudo-likelihood method.
复制标题

DOI:
10.1111/j.1541-0420.2010.01469.x
复制
发表时间:
2011-06
期刊:
影响因子:
1.9
通讯作者:
Begg CB
Begg CB
中科院分区:
数学3区
文献类型:
--
作者:
Capanu M;Begg CB

文献摘要

参考文献

被引文献

相似文献

许多主要基因已经被识别出来,它们强烈影响癌症的风险。然而,基因中通常会出现许多不同的突变,每一种突变都可能会增加风险,也可能不会。关键是要确定哪些特定突变是有害的,哪些是无害的,以便从基因检测中了解到自己有突变的个人可以得到适当的咨询。这是一项具有挑战性的任务,因为不断发现新的突变,而关于每个突变的证据通常相对较少。在之前的一篇文章中,我们使用了使用伪似然和Gibbs抽样方法的分层建模,使用病例对照研究的数据来估计单个罕见变异的相对风险,并表明可以从分层模型的聚合能力中获得力量,以区分导致癌症风险的变异。然而,需要进一步的研究来验证渐近方法在这种稀疏数据中的应用。在这篇文章中,我们使用模拟来详细地研究伪似然方法的性质。我们还探索了两种可供选择的方法:对方差分量估计进行修正的伪似然方法和对方差分量进行贝叶斯估计的混合伪似然方法。我们通过观察估计量的偏差和覆盖性质以及分层建模估计相对于最大似然估计的效率来研究这些分层建模技术的有效性。结果表明,非常稀疏变量的相对风险估计具有较小的偏差,估计的95%可信区间通常是反保守的,尽管实际覆盖率通常在90%以上。随着第二阶段模型中残差方差的减小,置信度区间的宽度变窄。结果还表明,与传统的Logistic回归估计相比,分层建模估计具有更短的可信区间,并且这些相对改善随着变量变得越来越少而增加。
Many major genes have been identified that strongly influence the risk of cancer. However, there are typically many different mutations that can occur in the gene, each of which may or may not confer increased risk. It is critical to identify which specific mutations are harmful, and which ones are harmless, so that individuals who learn from genetic testing that they have a mutation can be appropriately counseled. This is a challenging task, since new mutations are continually being identified, and there is typically relatively little evidence available about each individual mutation. In an earlier article we employed hierarchical modeling using the pseudo-likelihood and Gibbs sampling methods to estimate the relative risks of individual rare variants using data from a case-control study and showed that one can draw strength from the aggregating power of hierarchical models to distinguish the variants that contribute to cancer risk. However, further research is needed to validate the application of asymptotic methods to such sparse data. In this article we use simulations to study in detail the properties of the pseudo-likelihood method for this purpose. We also explore two alternative approaches: pseudo-likelihood with correction for the variance component estimate as proposed by and a hybrid pseudo-likelihood approach with Bayesian estimation of the variance component. We investigate the validity of these hierarchical modeling techniques by looking at the bias and coverage properties of the estimators as well as at the efficiency of the hierarchical modeling estimates relative to that of the maximum likelihood estimates. The results indicate that the estimates of the relative risks of very sparse variants have small bias, and that the estimated 95% confidence intervals are typically anti-conservative, though the actual coverage rates are generally above 90 per cent. The widths of the confidence intervals narrow as the residual variance in the second-stage model is reduced. The results also show that the hierarchical modeling estimates have shorter confidence intervals relative to estimates obtained from conventional logistic regression, and that these relative improvements increase as the variants become more rare.
DOI: 10.1101/gr.176601
发表时间: 2001-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Ng, PC;Henikoff, S
通讯作者: Henikoff, S
DOI: 10.1038/ng.2007.53
发表时间: 2008-01-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Stratton, Michael R.;Rahman, Nazneen
通讯作者: Rahman, Nazneen
DOI: 10.1080/00949659308811554
发表时间: 1993-01-01
影响因子: 1.2
作者:
WOLFINGER, R;OCONNELL, M
通讯作者: OCONNELL, M
DOI: 10.1086/424388
发表时间: 2004-10-01
影响因子: 9.8
作者:
Goldgar, DE;Easton, DF;Couch, FJ
通讯作者: Couch, FJ
DOI: 10.1002/humu.21202
发表时间: 2010-03
期刊: HUMAN MUTATION
影响因子: 3.9
作者:
Borg, Ake;Haile, Robert W.;Malone, Kathleen E.;Capanu, Marinela;Diep, Ahn;Torngren, Therese;Teraoka, Sharon;Begg, Colin B.;Thomas, Duncan C.;Concannon, Patrick;Mellemkjaer, Lene;Bernstein, Leslie;Tellhed, Lina;Xue, Shanyan;Olson, Eric R.;Liang, Xiaolin;Dolle, Jessica;Borresen-Dale, Anne-Lise;Bernstein, Jonine L.
通讯作者: Bernstein, Jonine L.