GeneHancer: genome-wide integration of enhancers and target genes in GeneCards.

GeneHancer: genome-wide integration of enhancers and target genes in GeneCards.
复制标题

DOI:
10.1093/database/bax028
复制
发表时间:
2017-01-01
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Cohen D
Cohen D
中科院分区:
其他
文献类型:
--
作者:
Fishilevich S;Nudel R;Rappaport N;Hadar R;Plaschkes I;Iny Stein T;Rosen N;Kohn A;Twik M;Safran M;Lancet D;Cohen D

文献摘要

被引文献

相似文献

理解基因调控的一个主要挑战是明确识别增强子元件并揭示它们与基因的联系。我们目前的GeneHancer,一个新的数据库的人类增强子和他们推断的靶基因,在GeneCards的框架。首先,我们整合了来自四个不同的全基因组数据库的434000个已报告的增强子:DNA元件百科全书(ENCODE),Ensembl监管构建,哺乳动物基因组功能注释(FANTOM)项目和VISTA增强子浏览器。采用旨在消除冗余的整合算法,GeneHancer描绘了285 000个整合的候选增强子(覆盖基因组的12.4%),其中94 000个来自一个以上的来源,每个都分配了一个注释衍生的置信度得分。GeneHancer随后将增强子与基因联系起来,使用:基因和增强子RNA之间的组织共表达相关性,以及增强子靶向转录因子基因;增强子内变体的表达定量性状位点;以及捕获Hi-C,启动子特异性基因组构象测定。基于这四种方法中的每一种的个体得分,沿着基因-增强子基因组距离,形成GeneHancer的基于增强子-基因配对的组合似然性得分的基础。最后,我们定义了“精英”增强子-基因关系,反映了高似然增强子定义和强增强子-基因关联。GeneHancer预测完全集成在广泛使用的GeneCards Suite中,候选增强子及其注释显示在每个相关的GeneCard上。这有助于将非编码变体映射到增强子,并通过链接的基因,形成健康和疾病中全基因组序列的变体表型解释的基础。 数据库URL:http://www.genecards.org/
A major challenge in understanding gene regulation is the unequivocal identification of enhancer elements and uncovering their connections to genes. We present GeneHancer, a novel database of human enhancers and their inferred target genes, in the framework of GeneCards. First, we integrated a total of 434 000 reported enhancers from four different genome-wide databases: the Encyclopedia of DNA Elements (ENCODE), the Ensembl regulatory build, the functional annotation of the mammalian genome (FANTOM) project and the VISTA Enhancer Browser. Employing an integration algorithm that aims to remove redundancy, GeneHancer portrays 285 000 integrated candidate enhancers (covering 12.4% of the genome), 94 000 of which are derived from more than one source, and each assigned an annotation-derived confidence score. GeneHancer subsequently links enhancers to genes, using: tissue co-expression correlation between genes and enhancer RNAs, as well as enhancer-targeted transcription factor genes; expression quantitative trait loci for variants within enhancers; and capture Hi-C, a promoter-specific genome conformation assay. The individual scores based on each of these four methods, along with gene–enhancer genomic distances, form the basis for GeneHancer’s combinatorial likelihood-based scores for enhancer–gene pairing. Finally, we define ‘elite’ enhancer–gene relations reflecting both a high-likelihood enhancer definition and a strong enhancer–gene association. GeneHancer predictions are fully integrated in the widely used GeneCards Suite, whereby candidate enhancers and their annotations are displayed on every relevant GeneCard. This assists in the mapping of non-coding variants to enhancers, and via the linked genes, forms a basis for variant–phenotype interpretation of whole-genome sequences in health and disease. Database URL: http://www.genecards.org/