A novel genome-information content-based statistic for genome-wide association analysis designed for next-generation sequencing data.

A novel genome-information content-based statistic for genome-wide association analysis designed for next-generation sequencing data.
复制标题

一种新颖的基于基因组信息内容的统计数据,用于专为下一代测序数据设计的全基因组关联分析。

DOI:
10.1089/cmb.2012.0035
复制
发表时间:
2012
期刊:
Journal of computational biology : a journal of computational molecular cell biology
影响因子:
--
通讯作者:
Xiong,Momiao
Xiong,Momiao
中科院分区:
--
文献类型:
--
作者:
Luo,Li;Zhu,Yun;Xiong,Momiao

文献摘要

相似文献

为下一代测序数据设计的全基因组关联研究(GWAS)涉及测试基因组变异的关联,包括常见、低频率和罕见变异。目前的关联研究策略已被很好地开发,用于识别常见变异与常见疾病的关联,但当序列数据中存在大量等位基因异质性时,可能不适合。最近,组测试,分析他们的集体之间的频率差异的情况下,控制转移到罕见变异的关联分析中的多个变异的GWAS的常见变异的变异分析范式。然而,群体检验忽略了不同基因组位置的SNP之间的遗传效应差异。作为组检验的替代方法,我们开发了一种新的基于基因组信息内容的统计方法,用于检验基因组变异的整个等位基因频谱与疾病的关联。为了评估所提出的统计量的性能,我们使用基于1000个基因组计划中的全基因组低覆盖试点数据的大规模模拟来计算7种替代统计量的1型错误率和功效:基于基因组信息含量的统计量、广义T2、塌陷方法、多变量塌陷(CMC)方法、个体χ 2检验、加权和统计量、和可变阈值统计。最后,我们将这七个统计量应用于达拉斯心脏研究中已发表的ANGPTL3、ANGPTL4、ANGPTL5和ANGPTL6基因重测序数据集。我们报告说,在模拟和经验数据集中,基于基因组信息内容的统计量显着改善了第1类错误率,并且比其他六种统计量具有更高的功效。
The genome-wide association studies (GWAS) designed for next-generation sequencing data involve testing association of genomic variants, including common, low frequency, and rare variants. The current strategies for association studies are well developed for identifying association of common variants with the common diseases, but may be ill-suited when large amounts of allelic heterogeneity are present in sequence data. Recently, group tests that analyze their collective frequency differences between cases and controls shift the current variant-by-variant analysis paradigm for GWAS of common variants to the collective test of multiple variants in the association analysis of rare variants. However, group tests ignore differences in genetic effects among SNPs at different genomic locations. As an alternative to group tests, we developed a novel genome-information content-based statistics for testing association of the entire allele frequency spectrum of genomic variation with the diseases. To evaluate the performance of the proposed statistics, we use large-scale simulations based on whole genome low coverage pilot data in the 1000 Genomes Project to calculate the type 1 error rates and power of seven alternative statistics: a genome-information content-based statistic, the generalized T2, collapsing method, multivariate and collapsing (CMC) method, individualχ2test, weighted-sum statistic, and variable threshold statistic. Finally, we apply the seven statistics to published resequencing dataset fromANGPTL3, ANGPTL4,ANGPTL5,andANGPTL6genes in the Dallas Heart Study. We report that the genome-information content-based statistic has significantly improved type 1 error rates and higher power than the other six statistics in both simulated and empirical datasets.