The value of statistical or bioinformatics annotation for rare variant association with quantitative trait.

The value of statistical or bioinformatics annotation for rare variant association with quantitative trait.
复制标题

DOI:
10.1002/gepi.21747
复制
发表时间:
2013-11
影响因子:
2.1
通讯作者:
Li, Yun
Li, Yun
中科院分区:
医学4区
文献类型:
--
作者:
Byrnes, Andrea E.;Wu, Michael C.;Wright, Fred A.;Li, Mingyao;Li, Yun

文献摘要

参考文献

被引文献

相似文献

在过去的几年里,已经提出了大量的方法,罕见的变异与表型的关联。这些方法聚集了来自基因组区域的多个罕见变体的信息,但对于哪种方法最有效几乎没有共识。在汇总各种变量的信息时采用的加权办法是有效性的主要决定因素之一。在这里,我们提出了一个系统的评价多个加权方案,通过一系列的模拟旨在模仿大型测序研究的数量性状。我们评估现有的表型独立和依赖的方法,以及惩罚回归方法,包括拉索,弹性网络和SCAD估计的权重。我们发现,当高质量的功能注释可用时,表型依赖计划之间的功率差异可以忽略不计。当功能注释不可用或不完整时,所有方法都遭受功率损失;然而,变量选择方法以增加计算时间为代价胜过其他方法。因此,在缺乏良好注释的情况下,我们建议对表型独立加权方案所涉及的顶部区域采用变量选择方法(可视为“统计注释”)。此外,一旦涉及到一个区域,变量选择可以帮助识别潜在的因果SNP用于生物学验证。这些发现得到了对1898名个体的高覆盖率靶向测序研究的分析的支持。
In the past few years, a plethora of methods for rare variant association with phenotype have been proposed. These methods aggregate information from multiple rare variants across genomic region(s), but there is little consensus as to which method is most effective. The weighting scheme adopted when aggregating information across variants is one of the primary determinants of effectiveness. Here we present a systematic evaluation of multiple weighting schemes through a series of simulations intended to mimic large sequencing studies of a quantitative trait. We evaluate existing phenotype-independent and -dependent methods, as well as weights estimated by penalized regression approaches including Lasso, Elastic Net and SCAD. We find that the difference in power between phenotype-dependent schemes is negligible when high quality functional annotations are available. When functional annotations are unavailable or incomplete, all methods suffer from power loss; however, the variable selection methods outperform the others at the cost of increased computational time. Therefore, in the absence of good annotation, we recommend variable selection methods (which can be viewed as “statistical annotation”) on top regions implicated by a phenotype independent weighting scheme. Further, once a region is implicated, variable selection can help to identify potential causal SNPs for biological validation. These findings are supported by an analysis of a high coverage targeted sequencing study of 1898 individuals.
来自1,092个人基因组的遗传变异的综合图。
DOI: 10.1038/nature11632
发表时间: 2012-11-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
COLAUS研究:一项基于人群的研究,旨在研究心血管危险因素和代谢综合征的流行病学和遗传决定因素。
DOI: 10.1186/1471-2261-8-6
发表时间: 2008-03-17
影响因子: 2.1
作者:
Firmann, Mathieu;Mayor, Vladimir;Vidal, Pedro Marques;Bochud, Murielle;Pecoud, Alain;Hayoz, Daniel;Paccaud, Fred;Preisig, Martin;Song, Kijoung S.;Yuan, Xin;Danoff, Theodore M.;Stirnadel, Heide A.;Waterworth, Dawn;Mooser, Vincent;Waeber, Gerard;Vollenweider, Peter
通讯作者: Vollenweider, Peter
DOI: 10.1073/pnas.0812824106
发表时间: 2009-03-10
影响因子: 11.1
作者:
Kryukov, Gregory V.;Shpunt, Alexander;Sunyaev, Shamil R.
通讯作者: Sunyaev, Shamil R.
DOI: 10.1016/j.ajhg.2012.08.031
发表时间: 2012-11-02
影响因子: 9.8
作者:
Auer, Paul L.;Johnsen, Jill M.;Li, Yun
通讯作者: Li, Yun
DOI: 10.1016/j.ajhg.2008.06.024
发表时间: 2008-09-12
影响因子: 9.8
作者:
Li, Bingshan;Leal, Suzanne M.
通讯作者: Leal, Suzanne M.