A novel genome-information content-based statistic for genome-wide association analysis designed for next-generation sequencing data.
A novel genome-information content-based statistic for genome-wide association analysis designed for next-generation sequencing data.
复制标题
一种新颖的基于基因组信息内容的统计数据,用于专为下一代测序数据设计的全基因组关联分析。
DOI:
10.1089/cmb.2012.0035
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Xiong,Momiao
中科院分区:
文献类型:
--
作者:
Luo,Li;Zhu,Yun;Xiong,Momiao
The genome-wide association studies (GWAS) designed for next-generation sequencing data involve testing association of genomic variants, including common, low frequency, and rare variants. The current strategies for association studies are well developed for identifying association of common variants with the common diseases, but may be ill-suited when large amounts of allelic heterogeneity are present in sequence data. Recently, group tests that analyze their collective frequency differences between cases and controls shift the current variant-by-variant analysis paradigm for GWAS of common variants to the collective test of multiple variants in the association analysis of rare variants. However, group tests ignore differences in genetic effects among SNPs at different genomic locations. As an alternative to group tests, we developed a novel genome-information content-based statistics for testing association of the entire allele frequency spectrum of genomic variation with the diseases. To evaluate the performance of the proposed statistics, we use large-scale simulations based on whole genome low coverage pilot data in the 1000 Genomes Project to calculate the type 1 error rates and power of seven alternative statistics: a genome-information content-based statistic, the generalized T2, collapsing method, multivariate and collapsing (CMC) method, individualχ2test, weighted-sum statistic, and variable threshold statistic. Finally, we apply the seven statistics to published resequencing dataset fromANGPTL3, ANGPTL4,ANGPTL5,andANGPTL6genes in the Dallas Heart Study. We report that the genome-information content-based statistic has significantly improved type 1 error rates and higher power than the other six statistics in both simulated and empirical datasets.