A novel adaptive method for the analysis of next-generation sequencing data to detect complex trait associations with rare variants due to gene main effects and interactions.

A novel adaptive method for the analysis of next-generation sequencing data to detect complex trait associations with rare variants due to gene main effects and interactions.
复制标题

DOI:
10.1371/journal.pgen.1001156
复制
发表时间:
2010-10-14
期刊:
影响因子:
4.5
通讯作者:
Leal SM
Leal SM
中科院分区:
生物学2区
文献类型:
--
作者:
Liu DJ;Leal SM

文献摘要

参考文献

被引文献

相似文献

有确凿的证据表明,罕见的变异有助于复杂的疾病病因。下一代测序技术使发现候选基因、外显子组和基因组中的罕见变异成为可能。在一个新的框架下,基于核的自适应聚类(KBAC)被开发用于执行强大的基于基因/位点的罕见变异关联测试。KBAC将变异分类和关联测试结合在一个连贯的框架中。协变量也可以纳入分析中,以控制潜在的混杂因素,包括年龄、性别和群体亚结构。为了评估KBAC的功效:1)使用欧洲人和非洲人的严格群体遗传模型模拟变异数据,其中参数从序列数据估计,以及2)使用由复杂疾病(包括乳腺癌和先天性巨结肠症)激发的模型生成表型。它表明,KBAC具有上级权力相比,其他罕见的变异分析方法,如组合的多元和崩溃和重量和统计。在存在变异误分类和基因相互作用的情况下,使用KBAC的关联测试是特别有利的。KBAC方法还应用于测试关联,使用来自达拉斯心脏研究的序列数据,能量代谢性状与ANGPTL 3、4、5和6基因中的罕见变体之间。一些新的协会被确定,包括高密度脂蛋白和极低密度脂蛋白与ANGPTL 4的协会。KBAC方法在一个用户友好的R包中实现。已经证明,罕见和常见的变异都涉及复杂的疾病病因。直到最近,才有可能对常见变异进行大规模分析。随着下一代测序技术的发展,罕见变异的检测和定位已经成为可能。然而,用于分析常见变异的方法对于分析罕见变异并不强大。为了解决罕见变异分析在一个新的框架下工作的问题,基于核的自适应聚类(KBAC)方法被开发来执行基于基因/位点的分析。KBAC将变异分类和关联测试结合在一个连贯的框架中。通过对群体遗传学和疾病数据的模拟,证明了KBAC具有上级能力,特别是在变异错误分类和基因相互作用的情况下。使用来自达拉斯心脏研究的数据,应用KBAC方法来检验能量代谢性状与ANGPTL 3、4、5和6基因中的罕见变异之间的关联。发现了一些新的关联。KBAC方法在一个用户友好的R包中实现。
There is solid evidence that rare variants contribute to complex disease etiology. Next-generation sequencing technologies make it possible to uncover rare variants within candidate genes, exomes, and genomes. Working in a novel framework, the kernel-based adaptive cluster (KBAC) was developed to perform powerful gene/locus based rare variant association testing. The KBAC combines variant classification and association testing in a coherent framework. Covariates can also be incorporated in the analysis to control for potential confounders including age, sex, and population substructure. To evaluate the power of KBAC: 1) variant data was simulated using rigorous population genetic models for both Europeans and Africans, with parameters estimated from sequence data, and 2) phenotypes were generated using models motivated by complex diseases including breast cancer and Hirschsprung's disease. It is demonstrated that the KBAC has superior power compared to other rare variant analysis methods, such as the combined multivariate and collapsing and weight sum statistic. In the presence of variant misclassification and gene interaction, association testing using KBAC is particularly advantageous. The KBAC method was also applied to test for associations, using sequence data from the Dallas Heart Study, between energy metabolism traits and rare variants in ANGPTL 3,4,5 and 6 genes. A number of novel associations were identified, including the associations of high density lipoprotein and very low density lipoprotein with ANGPTL4. The KBAC method is implemented in a user-friendly R package. It has been demonstrated that both rare and common variants are involved in complex disease etiology. Until recently it was only possible to perform large scale analysis of common variants. With the development of next-generation sequencing technologies, detection and mapping of rare variants have been made possible. However, methods used to analyze common variants are not powerful for the analysis of rare variants. To address the problems of rare variant analysis working in a novel framework, the kernel-based adaptive cluster (KBAC) method was developed to perform gene/locus based analysis. The KBAC combines variant classification and association testing in a coherent framework. Through simulations motivated by population genetic and disease data, it is demonstrated that the KBAC has superior power to other rare variant analysis methods, especially in the presence of variant misclassification and gene interaction. Using data from the Dallas Heart Study, the KBAC method was applied to test for associations between energy metabolism traits and rare variants in ANGPTL 3,4,5 and 6 genes. A number of novel associations were identified. The KBAC method is implemented in a user-friendly R package.
DOI: 10.1093/bioinformatics/btn522
发表时间: 2008-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hernandez, Ryan D.
通讯作者: Hernandez, Ryan D.
DOI: 10.1073/pnas.0812824106
发表时间: 2009-03-10
影响因子: 11.1
作者:
Kryukov, Gregory V.;Shpunt, Alexander;Sunyaev, Shamil R.
通讯作者: Sunyaev, Shamil R.
DOI: 10.1002/hep.20466
发表时间: 2004-12-01
期刊: HEPATOLOGY
影响因子: 13.5
作者:
Browning, JD;Szczepaniak, LS;Hobbs, HH
通讯作者: Hobbs, HH
DOI: 10.1016/j.ajhg.2008.06.024
发表时间: 2008-09-12
影响因子: 9.8
作者:
Li, Bingshan;Leal, Suzanne M.
通讯作者: Leal, Suzanne M.
DOI: 10.1016/j.ajhg.2007.09.006
发表时间: 2008-01-01
影响因子: 9.8
作者:
Gorlov, Ivan P.;Gorlova, Olga Y.;Amos, Christopher I.
通讯作者: Amos, Christopher I.