Association studies for next-generation sequencing

Association studies for next-generation sequencing
复制标题

DOI:
10.1101/gr.115998.110
复制
发表时间:
2011-07-01
期刊:
影响因子:
7
通讯作者:
Xiong, Momiao
Xiong, Momiao
中科院分区:
生物学1区
文献类型:
--
作者:
Luo, Li;Boerwinkle, Eric;Xiong, Momiao

文献摘要

被引文献

相似文献

全基因组关联研究(GWAS)已成为识别影响复杂疾病的常见变异基因的主要方法。尽管取得了相当大的进展,但GWAS确定的常见变异仅占疾病遗传力的一小部分,不太可能解释常见疾病的大多数表型变异。缺失遗传力的一个潜在来源是罕见变异的贡献。下一代测序技术将检测数百万种新型罕见变异,但这些技术有三个定义性特征:识别大量罕见变异、高比例的序列错误和大比例的缺失数据。这些特征对测试罕见变异与感兴趣的表型的关联提出了挑战。在这项研究中,我们使用基因组连续体模型和功能主成分作为开发新的和强大的关联分析方法的一般原则,设计用于重测序数据。我们使用模拟来计算I类错误率和九种替代统计量的功效:基于两个函数主成分分析(FPCA)的统计量,基于多变量主成分分析(MPCA)的统计量,加权和(WSS),可变阈值(VT)方法,广义T-2,折叠方法,CMC方法和个体卡方检验。我们还研究了序列错误对其I类错误率的影响。最后,我们将9个统计量应用于达拉斯心脏研究中ANGPTL 4的已发表重测序数据集。我们报告说,基于FPCA的统计有更高的权力,以检测关联的罕见变异和更强的能力,过滤序列错误比其他七种方法。
Genome-wide association studies (GWAS) have become the primary approach for identifying genes with common variants influencing complex diseases. Despite considerable progress, the common variations identified by GWAS account for only a small fraction of disease heritability and are unlikely to explain the majority of phenotypic variations of common diseases. A potential source of the missing heritability is the contribution of rare variants. Next-generation sequencing technologies will detect millions of novel rare variants, but these technologies have three defining features: identification of a large number of rare variants, a high proportion of sequence errors, and a large proportion of missing data. These features raise challenges for testing the association of rare variants with phenotypes of interest. In this study, we use a genome continuum model and functional principal components as a general principle for developing novel and powerful association analysis methods designed for resequencing data. We use simulations to calculate the type I error rates and the power of nine alternative statistics: two functional principal component analysis (FPCA)-based statistics, the multivariate principal component analysis (MPCA)-based statistic, the weighted sum (WSS), the variable-threshold (VT) method, the generalized T-2, the collapsing method, the CMC method, and individual chi(2) tests. We also examined the impact of sequence errors on their type I error rates. Finally, we apply the nine statistics to the published resequencing data set from ANGPTL4 in the Dallas Heart Study. We report that FPCA-based statistics have a higher power to detect association of rare variants and a stronger ability to filter sequence errors than the other seven methods.