Extending Rare-Variant Testing Strategies Analysis of Noncoding Sequence and Imputed Genotypes

Extending Rare-Variant Testing Strategies Analysis of Noncoding Sequence and Imputed Genotypes
复制标题

DOI:
10.1016/j.ajhg.2010.10.012
复制
发表时间:
2010-11-12
影响因子:
9.8
通讯作者:
Zollner, Sebastian
Zollner, Sebastian
中科院分区:
生物学1区
文献类型:
--
作者:
Zawistowski, Matthew;Gopalakrishnan, Shyam;Zollner, Sebastian

文献摘要

被引文献

相似文献

下一代测序技术已经彻底改变了我们研究罕见遗传变异对可遗传性状的贡献的能力。然而,现有的单标记关联测试对于检测罕见风险变异的能力不足。更强大的方法涉及将来自同一基因的多个罕见变异联合收割机组合成单个测试统计量的汇集方法。在本文中,我们考虑了一种直观且计算效率高的合并统计量累积次要等位基因检验(CMAT)。我们评估了CMAT和其他合并方法在用群体遗传模型模拟的数据集上的性能,以包含现实水平的中性变异。我们考虑了从仅外显子到包含非编码变异的全基因分析的研究设计。设计CMAT达到功率与以前提出的方法,我们然后扩展CMAT的概率基因型,并描述应用低覆盖率测序和插补数据,我们表明,增加序列数据与插补样本是一个实用的方法,以增加罕见变异研究的权力,我们还提供了一种方法,控制混杂变量,如人口分层,最后,我们证明,我们的方法使得使用外部插补模板来分析插补到现有GWAS数据集中的罕见变异成为可能。作为原理的证明,我们对超过800万个SNP进行了CMAT分析,这些SNP通过使用来自1000个基因组计划的单倍型插补到GAIN银屑病数据集中。
Next Generation Sequencing Technology has revolutionized our ability to study the contribution of rare genetic variation to heritable traits However existing single marker association tests are underpowered for detecting rare risk variants A more powerful approach involves pooling methods that combine multiple rare variants from the same gene into a single test statistic Proposed pooling methods can be limited because they generally assume high quality genotypes derived from deep coverage sequencing which may not be avail able In this paper we consider an intuitive and computationally efficient pooling statistic the cumulative minor allele test (CMAT) We assess the performance of the CMAT and other pooling methods on datasets simulated with population genetic models to contain realistic levels of neutral variation We consider study designs ranging from exon only to whole gene analyses that contain noncoding variants For all study designs the CMAT achieves power comparable to that of previously proposed methods We then extend the CMAT to probabilistic genotypes and describe application to low coverage sequencing and imputation data We show that augmenting sequence data with imputed samples is a practical method for increasing the power of rare variant studies We also provide a method of controlling for confounding variables such as population stratification Finally we demonstrate that our method makes it possible to use external imputation templates to analyze rare variants imputed into existing GWAS datasets As proof of principle, we performed a CMAT analysis of more than 8 million SNPs that we imputed Into the GAIN psoriasis dataset by using haplotypes from the 1000 Genomes Project