The Intolerance of Regulatory Sequence to Genetic Variation Predicts Gene Dosage Sensitivity.

The Intolerance of Regulatory Sequence to Genetic Variation Predicts Gene Dosage Sensitivity.
复制标题

DOI:
10.1371/journal.pgen.1005492
复制
发表时间:
2015-09
期刊:
影响因子:
4.5
通讯作者:
Goldstein DB
Goldstein DB
中科院分区:
生物学2区
文献类型:
--
作者:
Petrovski S;Gussow AB;Wang Q;Halvorsen M;Han Y;Weir WH;Allen AS;Goldstein DB

文献摘要

被引文献

相似文献

非编码序列含有致病突变。然而,与蛋白质编码序列中的突变相比,致病性调控突变是出了名的难以识别。最根本的是,我们还不善于识别人类基因组中对调节基因表达最重要的序列片段。由于这个原因,很难对调控区应用同样类型的分析范式,这些范式已成功地应用于识别影响风险的蛋白质编码区中的突变。为了确定剂量敏感基因是否在其非编码序列中具有不同的模式,我们提出了两种主要的方法,只专注于基因的近端非编码调控序列。第一种方法是最近引入的残差变异不耐受评分(RVIS)的调控序列类似物,称为非编码RVIS或ncRVIS。ncRVIS比较了人类基因调控序列中观察到的和预测的持续变异水平。第二种方法称为ncGERP,反映了使用GERP++的基因调控序列的系统发育保守性。我们评估了这两种方法与四种基因列表的相关性,这些基因列表使用不同的方法来识别已知或可能通过表达变化引起疾病的基因:1)已知通过单倍不足引起疾病的基因,2)在ClinGen的基因组剂量图中被策展为剂量敏感的基因,3)被判断可能处于改变表达水平的突变的纯化选择下的基因,因为它们在一般群体中统计学上耗尽了功能丧失变体,和4)基于在一般群体中存在拷贝数变异体而判断为不太可能引起疾病的基因。我们发现,这两个非编码得分是高度预测剂量敏感性使用这些标准中的任何一个。以类似的方式ncGERP,我们评估两个整体为基础的预测区域非编码的重要性,ncCADD和ncGWAVA,并发现这两个分数是显着预测人类剂量敏感基因,并出现携带信息超出保护,评估ncGERP。这些结果强调了人类基因组中非编码序列延伸的不耐受性可以为其他基因组注释方法提供关键的补充工具,以帮助识别人类基因组中越来越可能含有影响疾病风险的突变的部分。非编码序列的突变可能会导致疾病,但很难识别。在这里,我们提出了两种方法,旨在帮助确定基因组的非编码区,可能携带突变影响疾病。第一种方法是基于比较观察到的和预测水平的常设人类变异的非编码外显子序列的基因。第二种方法是基于使用GERP++的基因的非编码外显子序列的系统发育保守性。我们发现,这两种方法都可以通过表达水平的变化预测已知引起疾病的基因,一般人群中功能丧失等位基因的基因,以及一般人群中允许拷贝数变异的基因。我们发现,这两个分数有助于解释功能丧失突变,并在定义区域的非编码序列,更有可能窝藏突变,影响疾病的风险。
Noncoding sequence contains pathogenic mutations. Yet, compared with mutations in protein-coding sequence, pathogenic regulatory mutations are notoriously difficult to recognize. Most fundamentally, we are not yet adept at recognizing the sequence stretches in the human genome that are most important in regulating the expression of genes. For this reason, it is difficult to apply to the regulatory regions the same kinds of analytical paradigms that are being successfully applied to identify mutations among protein-coding regions that influence risk. To determine whether dosage sensitive genes have distinct patterns among their noncoding sequence, we present two primary approaches that focus solely on a gene’s proximal noncoding regulatory sequence. The first approach is a regulatory sequence analogue of the recently introduced residual variation intolerance score (RVIS), termed noncoding RVIS, or ncRVIS. The ncRVIS compares observed and predicted levels of standing variation in the regulatory sequence of human genes. The second approach, termed ncGERP, reflects the phylogenetic conservation of a gene’s regulatory sequence using GERP++. We assess how well these two approaches correlate with four gene lists that use different ways to identify genes known or likely to cause disease through changes in expression: 1) genes that are known to cause disease through haploinsufficiency, 2) genes curated as dosage sensitive in ClinGen’s Genome Dosage Map, 3) genes judged likely to be under purifying selection for mutations that change expression levels because they are statistically depleted of loss-of-function variants in the general population, and 4) genes judged unlikely to cause disease based on the presence of copy number variants in the general population. We find that both noncoding scores are highly predictive of dosage sensitivity using any of these criteria. In a similar way to ncGERP, we assess two ensemble-based predictors of regional noncoding importance, ncCADD and ncGWAVA, and find both scores are significantly predictive of human dosage sensitive genes and appear to carry information beyond conservation, as assessed by ncGERP. These results highlight that the intolerance of noncoding sequence stretches in the human genome can provide a critical complementary tool to other genome annotation approaches to help identify the parts of the human genome increasingly likely to harbor mutations that influence risk of disease. Mutations in noncoding sequence can cause disease but are very difficult to recognize. Here, we present two approaches intended to help identify noncoding regions of the genome that may carry mutations influencing disease. The first approach is based on comparing observed and predicted levels of standing human variation in the noncoding exome sequence of a gene. The second approach is based on the phylogenetic conservation of a gene’s noncoding exome sequence using GERP++. We find that both approaches can predict genes known to cause disease through changes in expression level, genes depleted of loss-of-function alleles in the general population, and genes permissive of copy number variants in the general population. We find that both scores aid in interpreting loss-of-function mutations and in defining regions of noncoding sequence that are more likely to harbor mutations that influence risk of disease.