AllergyGenDB: A literature and functional annotation-based omics database for allergic diseases.

AllergyGenDB: A literature and functional annotation-based omics database for allergic diseases.
复制标题

AllergyGenDB:基于文献和功能注释的过敏性疾病组学数据库。

DOI:
10.1111/all.14219
复制
发表时间:
2020
期刊:
影响因子:
12.4
通讯作者:
Mersha,TesfayeB
Mersha,TesfayeB
中科院分区:
医学1区
文献类型:
--
作者:
Chen,Siqi;Ghandikota,Sudhir;Gautam,Yadu;Mersha,TesfayeB

文献摘要

相似文献

在包括测序、基因分型和大数据分析在内的高通量技术最近爆炸性增长的推动下,针对过敏性疾病的大量多组学和功能注释数据正在积累。然而,这些资源分散在不同格式的各种数据库中,手动搜索与过敏性疾病相关的基因是一项艰巨的任务。随着生物医学文献的积累超过了大多数研究人员和临床医生与其所在领域保持同步的能力,文本挖掘至关重要。文本挖掘涉及通过搜索文档中的文本字符串并分析其频率和上下文来自动提取信息。因此,需要一个全面的基于网络的一站式生物信息学工具来有效地查询和检索研究结果,并得出新的假设,以了解变态反应性疾病在基因、变异或途径水平上的独特和共享的关联(或共同发生)。在这里,我们描述AllergyGenDB系统地从2800万PubMed生物医学文献引文中产生与过敏相关的基因(或变体),并从93,892个GWASCatalog和DBGaP SNPs资源1、2中提取它们的功能注释信息,从Roadmap表观基因组学、GTEx和ENCODE数据中提取。这实现了三件事。首先,它将通过证明已知基因如预测的那样发生来验证实验。其次,它将迅速强调哪些基因得到了文献的支持,哪些基因在给定的背景下是新的。第三,它将引导新的假设,在基因、变异或途径水平上理解变态反应性疾病的独特和共同的联系(或共同发生)。
Driven mostly by the recent explosion of high-throughput technologies including sequencing, genotyping, and big data analytics, a vast amount of multi-omics and functional annotation data are accumulating for allergic diseases. Yet, these resources are scattered across various databases with different formats, and conducting manual search for genes associated with allergic diseases is a formidable task. As the accumulation of biomedical literature outpaces the ability of most researchers and clinicians to stay abreast of their fields, text mining, which involves automated information extraction by searching documents for text strings and analyzing their frequency and context, is critical. Hence, a comprehensive web-based one-stop bioinformatics tool is needed to efficiently query and retrieve research findings and draw novel hypotheses to understand the unique and shared association (or co-occurrence) of allergic diseases at gene, variant or pathway level. Here we describe AllergyGenDB to systematically generated allergy associated genes (or variants) from over 28 million PubMed biomedical literature citation, and from 93,892 GWAS Catalog and dbGaP SNPs resources, 1, 2 and distill their functional annotation information from Roadmap Epigenomics, GTEx, and ENCODE data. This accomplishes three things. First, it would serve to validate experiments by demonstrating that known genes occur as predicted. Second, it would rapidly highlight which genes are supported by the literature and which genes are novel in a given context. Third, it would lead novel hypotheses for understanding the unique and shared association (or co-occurrence) of allergic diseases at gene, variant or pathway level.