Rare-Variant Association Testing for Sequencing Data with the Sequence Kernel Association Test

Rare-Variant Association Testing for Sequencing Data with the Sequence Kernel Association Test
复制标题

DOI:
10.1016/j.ajhg.2011.05.029
复制
发表时间:
2011-07-15
影响因子:
9.8
通讯作者:
Lin, Xihong
Lin, Xihong
中科院分区:
生物学1区
文献类型:
--
作者:
Wu, Michael C.;Lee, Seunggeun;Lin, Xihong

文献摘要

被引文献

相似文献

进行测序研究越来越多地识别与复杂性状相关的稀有变体。在此类研究中,对稀有变体的经典单标记关联分析的有限能力构成了核心挑战。我们提出了序列内核关联测试(SKAT),这是一种受监督的,灵活的计算有效回归方法,以测试一个区域中遗传变异(常见和稀有)与连续或二分法性状之间的关联,同时易于调整协变量。作为基于分数的方差 - 组件测试,SKAT可以通过仅拟合包含协变量的空模型来迅速计算P值,因此可以轻松地应用于全基因组数据。使用SKAT分析1000个个体的全基因组测序研究,通过将整个基因组分割为30 Kb区域,仅需在笔记本电脑上7小时。通过分析达拉斯心脏研究中广泛的实用场景和甘油三酸酯数据的模拟数据,我们表明,SKAT可以大大优于几个替代较差的稀有关联测试。我们还提供分析能力和样品大小的计算,以帮助设计候选基因,全基因组和全基因组序列关联研究。
Sequencing studies are increasingly being conducted to identify rare variants associated with complex traits. The limited power of classical single-marker association analysis for rare variants poses a central challenge in such studies. We propose the sequence kernel association test (SKAT), a supervised, flexible, computationally efficient regression method to test for association between genetic variants (common and rare) in a region and a continuous or dichotomous trait while easily adjusting for covariates. As a score-based variance-component test, SKAT can quickly calculate p values analytically by fitting the null model containing only the covariates, and so can easily be applied to genome-wide data. Using SKAT to analyze a genome-wide sequencing study of 1000 individuals, by segmenting the whole genome into 30 kb regions, requires only 7 hr on a laptop. Through analysis of simulated data across a wide range of practical scenarios and triglyceride data from the Dallas Heart Study, we show that SKAT can substantially outperform several alternative rare-variant association tests. We also provide analytic power and sample-size calculations to help design candidate-gene, whole-exome, and whole-genome sequence association studies.