Comprehensive repository of regulatory genomic features and their role in human disease
Comprehensive repository of regulatory genomic features and their role in human disease
批准号:
426114092
负责人:
Professor Dr. Ulf Leser
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Units
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:
中文摘要
DNA调控元件的研究在生物医学研究中有着悠久的传统。存在大量的数据,从单个基因活动的孤立测量到功能研究和国际协调的全基因组调查。全面和高质量地概述关于人类基因调控的知识现状是规划未来实验的重要前提。然而,有针对性的、高质量的实验结果只在科学文章中发表。反过来,高通量实验收集的监管数据分散在大量数据库中。我们现在的目标是制定并向国际社会提供一份全面的目录,其中包括与人类疾病有关的这些区域的调控基因组特征和变异。我们的项目分为数据集成(DI)和信息提取(IE)部分。在上一个资助期间,我们开发了第一个用监管信息注释的文本语料库。这被用来训练文本挖掘算法,这些算法可以检测新文本中的调控序列元素。这导致了这些元素及其与基因和疾病的假定关联的第一个文本挖掘派生的集合。此外,我们还提出了一种基于深度神经网络和大型语言模型的实体归一化方法。在第二个应用阶段,对于DI,我们将专注于更新和扩大集成数据库的数量,并使集成过程自动化。在IE中,我们将把重点从实体识别和规范化转移到实体关系提取上。训练模型将需要扩展语料库的注释,以表示调控特征与基因、变异和疾病之间的关系。我们将在这个扩展的语料库上训练最先进的关系提取方法,并将训练的模型应用于特定疾病的文本集合。为了管理结果,我们将开发一种创新的快速注释方法,重点放在用户满意度和易用性上,这一方面在目前可用的软件工具中仍未得到充分研究。为了快速方便地访问项目中集成、提取和管理的所有监管特征数据,我们将开发一个具有直观可视化的用户友好的网络界面,并将其集成到RegulationSpotter网站中。
英文摘要
The study of regulatory DNA elements has a long tradition in biomedical research. A flood of data exists, ranging from isolated measurements of single gene activities to functional studies and internationally coordinated genome-wide investigations. A comprehensive and high-quality overview of the current state of knowledge regarding gene regulation in humans is an important prerequisite for planning future experiments. However, the results of targeted, high-quality experiments are only published in scientific articles. In turn, regulatory data collected by high-throughput experiments are scattered across a large number of databases. We now aim to develop and make available to the international community a comprehensive catalog of regulatory genomic features and variation in these regions that relate to human diseases. Our project is divided into a data integration (DI) and an information extraction (IE) part. In the last funding period, we developed the first text corpus annotated with regulatory information. This was used to train text-mining algorithms that can detect regulatory sequence elements in new texts. This resulted in the first text-mining-derived collection of these elements and their putative associations with genes and diseases. In addition, we have developed an entity normalization method based on Deep Neural Networks and large language models. In the second application phase, for DI, we will focus on updating and expanding the number of integrated databases and automating the integration process. In IE, we will shift our focus from entity recognition and normalization to entity relationship extraction. Training the models will require extending the annotation of the corpus to represent relationships between regulatory features and genes, variants, and diseases. We will train state-of-the-art methods for relation extraction on this extended corpus and apply the trained models to disease-specific text collections. For curating the results, we will develop an innovative method for rapid annotation that focuses on user satisfaction and ease of use, an aspect that is still understudied in currently available software tools. For quick and easy access to all regulatory feature data integrated, extracted, and curated in the project, we will develop a user-friendly web interface with intuitive visualizations that will be integrated into the RegulationSpotter website.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Learning Table Similarity Measures
-
批准号:388146305
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2017
-
负责人:Professor Dr. Ulf Leser
-
依托单位:
Web Data Analytics and Scientific Workflows
-
批准号:248359967
-
项目类别:Research Units
-
资助金额:$0.0万
-
财政年份:2013
-
负责人:Professor Dr. Ulf Leser
-
依托单位:
Scalable Information Extraction in Stratosphere
-
批准号:174466407
-
项目类别:Research Units
-
资助金额:$0.0万
-
财政年份:2010
-
负责人:Professor Dr. Ulf Leser
-
依托单位:
海外基金