Which Genetics Variants in DNase-Seq Footprints Are More Likely to Alter Binding?

Which Genetics Variants in DNase-Seq Footprints Are More Likely to Alter Binding?
复制标题

DOI:
10.1371/journal.pgen.1005875
复制
发表时间:
2016-02
期刊:
影响因子:
4.5
通讯作者:
Pique-Regi R
Pique-Regi R
中科院分区:
生物学2区
文献类型:
--
作者:
Moyerbrailean GA;Kalita CA;Harvey CT;Wen X;Luca F;Pique-Regi R

文献摘要

被引文献

相似文献

大量的实验工作正在表征调控基因组,但我们仍然缺乏对非编码区功能和沉默遗传变体的系统定义。在这里,我们将DNaseI足迹数据与基于序列的转录因子(TF)基序模型相结合,以预测遗传变异对153个组织和1,372个TF基序的TF结合的影响。我们得出的每个注释都是特定于细胞类型条件或分析的,并且是局部基序驱动的。我们在足迹中发现了580万个遗传变异,根据我们的模型预测,其中66%会影响TF结合。使用等位基因特异性高敏感性(ASH)的综合检查显示,只有后一组持续显示ASH的证据(3217个SNP,在20%FDR),这表明在足迹调节区的大多数(97%)遗传变异确实是沉默的。将这些信息与GWAS数据相结合,揭示了我们的注释有助于在GWAS命中区域对86个SNP进行计算精细映射,从而使挑选因果SNP的后验几率至少增加了2倍。受影响的组织特异性和假定的转铁蛋白结合位点的同一性提供的丰富的元信息也有助于识别支持这种关联的潜在机制。例如,在预测会影响HNF4结合位点的SNP中,与已经包括组织特异性注释的背景模型相比,与低密度脂蛋白水平相关的SNPs的富集率高出9.1倍。与复杂性状相关的遗传变异中,有很大一部分是在蛋白质编码基因之外发现的,可能会影响基因调控。许多实验工作都致力于绘制基因组中的调节区,但没有很多系统的方法来整合功能数据和调节序列来预测任何遗传变异对任何给定组织和基序的潜在影响。在这里,我们提出了一个特定于组织和因子的注释,它为常见和罕见的遗传变异提供了一个预测的功能效应。这些预测表明,基因调控区域中的大多数基因变异实际上是沉默的,其中一些预测得到了实验验证。对那些不是沉默的序列进行注释,可以让我们研究许多常见性状遗传结构的分子基础,也可以研究不同类型的调控序列在组织或转录因子之间的进化特性。总体而言,我们的研究支持这样的概念,即不同类别转录因子结合位点的多基因变异一直是导致人类疾病风险和复杂特征变异的进化力量的主要靶点。
Large experimental efforts are characterizing the regulatory genome, yet we are still missing a systematic definition of functional and silent genetic variants in non-coding regions. Here, we integrated DNaseI footprinting data with sequence-based transcription factor (TF) motif models to predict the impact of a genetic variant on TF binding across 153 tissues and 1,372 TF motifs. Each annotation we derived is specific for a cell-type condition or assay and is locally motif-driven. We found 5.8 million genetic variants in footprints, 66% of which are predicted by our model to affect TF binding. Comprehensive examination using allele-specific hypersensitivity (ASH) reveals that only the latter group consistently shows evidence for ASH (3,217 SNPs at 20% FDR), suggesting that most (97%) genetic variants in footprinted regulatory regions are indeed silent. Combining this information with GWAS data reveals that our annotation helps in computationally fine-mapping 86 SNPs in GWAS hit regions with at least a 2-fold increase in the posterior odds of picking the causal SNP. The rich meta information provided by the tissue-specificity and the identity of the putative TF binding site being affected also helps in identifying the underlying mechanism supporting the association. As an example, the enrichment for LDL level-associated SNPs is 9.1-fold higher among SNPs predicted to affect HNF4 binding sites than in a background model already including tissue-specific annotation. A large fraction of genetic variants that have been associated with complex traits are found outside of protein coding genes and likely affect gene regulation. Many experimental efforts have been dedicated to mapping regulatory regions in the genome but there are not many systematic methods that integrate functional data and regulatory sequences to predict the potential effect of any genetic variant on any given tissue and motif. Here we present a tissue and factor specific annotation that provides a predicted functional effect for both common and rare genetic variants. These predictions, certain of which are validated experimentally, show that the majority of genetic variants in gene regulatory regions are actually silent. Annotating those that are not silent allows us to investigate the molecular basis for the genetic architecture of many common traits and also to study the evolutionary properties that different types of regulatory sequences have across tissues or transcription factors. Overall, our study supports the concept that polygenic variation in binding sites for distinct classes of transcription factors has been a major target of evolutionary forces contributing to disease risk and complex trait variation in humans.