A hybrid likelihood model for sequence-based disease association studies.

A hybrid likelihood model for sequence-based disease association studies.
复制标题

用于基于序列的疾病关联研究的混合可能性模型。

DOI:
10.1371/journal.pgen.1003224
复制
发表时间:
2013
期刊:
影响因子:
4.5
通讯作者:
Karchin R
Karchin R
中科院分区:
生物学2区
文献类型:
--
作者:
Chen YC;Carter H;Parla J;Kramer M;Goes FS;Pirooznia M;Zandi PP;McCombie WR;Potash JB;Karchin R

文献摘要

参考文献

被引文献

相似文献

In the past few years, case-control studies of common diseases have shifted their focus from single genes to whole exomes. New sequencing technologies now routinely detect hundreds of thousands of sequence variants in a single study, many of which are rare or even novel. The limitation of classical single-marker association analysis for rare variants has been a challenge in such studies. A new generation of statistical methods for case-control association studies has been developed to meet this challenge. A common approach to association analysis of rare variants is the burden-style collapsing methods to combine rare variant data within individuals across or within genes. Here, we propose a new hybrid likelihood model that combines a burden test with a test of the position distribution of variants. In extensive simulations and on empirical data from the Dallas Heart Study, the new model demonstrates consistently good power, in particular when applied to a gene set (e.g., multiple candidate genes with shared biological function or pathway), when rare variants cluster in key functional regions of a gene, and when protective variants are present. When applied to data from an ongoing sequencing study of bipolar disorder (191 cases, 107 controls), the model identifies seven gene sets with nominal p-values0.05, of which one MAPK signaling pathway (KEGG) reaches trend-level significance after correcting for multiple testing. Inexpensive, high-throughput sequencing has transformed the field of case-control association studies. For the first time, it may be possible to identify the genetic underpinnings of complex diseases, by sequencing the DNA of hundreds (even thousands) of cases and controls and comparing patterns of DNA sequence variation. However, complex diseases are likely to be caused by many variants, some of which are very rare. Taken one at a time, the association between variant and disease phenotype may not be detectable by current statistical methods. One strategy is to identify regions where important variants occur by “collapsing” variants into groups. Here, we present a new collapsing approach, capable of detecting subtle genetic differences between cases and controls. We show, in extensive simulations and using a benchmark set of genes involved in human triglyceride levels, that the approach is potentially more powerful than existing methods. We apply the new method to an ongoing sequencing study of bipolar cases and controls and identify a set of genes found in neuronal synapses, which may be implicated in bipolar disorder.
DOI: 10.1093/bioinformatics/btr260
发表时间: 2011-06-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Liberzon, Arthur;Subramanian, Aravind;Mesirov, Jill P.
通讯作者: Mesirov, Jill P.
DOI: 10.1371/journal.pgen.1001156
发表时间: 2010-10-14
期刊: PLoS genetics
影响因子: 4.5
作者:
Liu DJ;Leal SM
通讯作者: Leal SM
DOI: 10.1093/bioinformatics/btn522
发表时间: 2008-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hernandez, Ryan D.
通讯作者: Hernandez, Ryan D.
DOI: 10.1093/nar/gkr1055
发表时间: 2012-01
影响因子: 14.9
作者:
Dreszer TR;Karolchik D;Zweig AS;Hinrichs AS;Raney BJ;Kuhn RM;Meyer LR;Wong M;Sloan CA;Rosenbloom KR;Roe G;Rhead B;Pohl A;Malladi VS;Li CH;Learned K;Kirkup V;Hsu F;Harte RA;Guruvadoo L;Goldman M;Giardine BM;Fujita PA;Diekhans M;Cline MS;Clawson H;Barber GP;Haussler D;James Kent W
通讯作者: James Kent W
DOI: 10.1073/pnas.0812824106
发表时间: 2009-03-10
影响因子: 11.1
作者:
Kryukov, Gregory V.;Shpunt, Alexander;Sunyaev, Shamil R.
通讯作者: Sunyaev, Shamil R.