BLOCK-BASED BAYESIAN EPISTASIS ASSOCIATION MAPPING WITH APPLICATION TO WTCCC TYPE 1 DIABETES DATA.

BLOCK-BASED BAYESIAN EPISTASIS ASSOCIATION MAPPING WITH APPLICATION TO WTCCC TYPE 1 DIABETES DATA.
复制标题

DOI:
10.1214/11-aoas469
复制
发表时间:
2011-09-01
期刊:
The annals of applied statistics
影响因子:
--
通讯作者:
Liu JS
Liu JS
中科院分区:
其他
文献类型:
--
作者:
Zhang BY;Zhang J;Liu JS

文献摘要

被引文献

相似文献

基因组中多个基因之间的相互作用可能会导致许多复杂的人类疾病的风险。全基因组单核苷酸多态性(SNPs)数据收集了成千上万的SNP标记,从成千上万的个人在病例对照设计的承诺,阐明我们对这种相互作用的理解。然而,由于连锁不平衡(LD),附近的SNP高度相关,可能的相互作用的数量太大,无法进行详尽的评估。我们提出了一种新的贝叶斯方法,用于同时将SNP划分为LD块,并在与疾病相关的块内选择SNP,无论是单独还是与其他SNP交互。当应用于同质群体数据时,该方法给出了LD块边界的后验概率,这不仅导致SNP的准确块划分,而且还提供了划分不确定性的度量。当应用于关联映射的病例对照数据时,该方法隐式地过滤掉仅由LD与相同块内的疾病位点创建的SNP关联。模拟研究表明,这种方法在检测多位点关联方面比我们测试的其他方法(包括我们的方法)更强大。当应用于WTCCC 1型糖尿病数据时,该方法鉴定了许多先前已知的T1D相关基因,包括PTPN 22,CTLA 4,MHC和IL 2RA。该方法还揭示了一些有趣的双向关联,这些关联是单一SNP方法无法检测到的。大多数重要的关联位于MHC区域内。我们的分析表明,MHC SNPs在几个已知的重组热点上形成长距离联合关联。通过控制MHC II类区域的单倍型,我们确定了MHC I类(HLA-A,HLA-B)和III类区域(BAT 1)的额外关联。我们还观察到在延伸的MHC区域中的基因PRSS 16、ZNF 184与MHC II类基因之间的显著相互作用。该方法可广泛应用于离散协变量相关的分类问题。
Interactions among multiple genes across the genome may contribute to the risks of many complex human diseases. Whole-genome single nucleotide polymorphisms (SNPs) data collected for many thousands of SNP markers from thousands of individuals under the case–control design promise to shed light on our understanding of such interactions. However, nearby SNPs are highly correlated due to linkage disequilibrium (LD) and the number of possible interactions is too large for exhaustive evaluation. We propose a novel Bayesian method for simultaneously partitioning SNPs into LD-blocks and selecting SNPs within blocks that are associated with the disease, either individually or interactively with other SNPs. When applied to homogeneous population data, the method gives posterior probabilities for LD-block boundaries, which not only result in accurate block partitions of SNPs, but also provide measures of partition uncertainty. When applied to case–control data for association mapping, the method implicitly filters out SNP associations created merely by LD with disease loci within the same blocks. Simulation study showed that this approach is more powerful in detecting multi-locus associations than other methods we tested, including one of ours. When applied to the WTCCC type 1 diabetes data, the method identified many previously known T1D associated genes, including PTPN22, CTLA4, MHC, and IL2RA. The method also revealed some interesting two-way associations that are undetected by single SNP methods. Most of the significant associations are located within the MHC region. Our analysis showed that the MHC SNPs form long-distance joint associations over several known recombination hotspots. By controlling the haplotypes of the MHC class II region, we identified additional associations in both MHC class I (HLA-A, HLA-B) and class III regions (BAT1). We also observed significant interactions between genes PRSS16, ZNF184 in the extended MHC region and the MHC class II genes. The proposed method can be broadly applied to the classification problem with correlated discrete covariates.