Comprehensive repository of regulatory genomic features and their role in human disease
Comprehensive repository of regulatory genomic features and their role in human disease
批准号:
426114092
负责人:
Professor Dr. Ulf Leser
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Units
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:
中文摘要
对DNA调控元件的研究在生物医学研究中有着悠久的传统。从单个基因活动的孤立测量到功能研究和国际协调的全基因组调查,存在大量数据。对人类基因调控的知识现状进行全面和高质量的概述是规划未来实验的重要先决条件。然而,有针对性的高质量实验的结果只会发表在科学文章中。反过来,高通量实验收集的监管数据分散在大量数据库中。我们现在的目标是编制并向国际社会提供这些区域与人类疾病有关的调控基因组特征和变异的全面目录。我们的项目分为数据集成(DI)和信息提取(IE)两个部分。在上一个资助期间,我们开发了第一个带有监管信息注释的文本语料库。这被用于训练文本挖掘算法,该算法可以检测新文本中的规则序列元素。这导致了这些元素的第一个文本挖掘衍生的集合,以及它们与基因和疾病的假定关联。此外,我们开发了一种基于深度神经网络和大型语言模型的实体归一化方法。在第二个应用阶段,对于DI,我们将专注于更新和扩展集成数据库的数量,并自动化集成过程。在IE中,我们将把重点从实体识别和规范化转移到实体关系提取上。训练模型需要扩展语料库的注释,以表示调节特征与基因、变异和疾病之间的关系。我们将在这个扩展语料库上训练最先进的关系提取方法,并将训练好的模型应用于特定疾病的文本集合。为了整理结果,我们将开发一种创新的快速注释方法,重点关注用户满意度和易用性,这是目前可用的软件工具尚未充分研究的一个方面。为了快速方便地访问项目中集成、提取和整理的所有监管特征数据,我们将开发一个具有直观可视化的用户友好网络界面,该界面将集成到RegulationSpotter网站中。
英文摘要
The study of regulatory DNA elements has a long tradition in biomedical research. A flood of data exists, ranging from isolated measurements of single gene activities to functional studies and internationally coordinated genome-wide investigations. A comprehensive and high-quality overview of the current state of knowledge regarding gene regulation in humans is an important prerequisite for planning future experiments. However, the results of targeted, high-quality experiments are only published in scientific articles. In turn, regulatory data collected by high-throughput experiments are scattered across a large number of databases. We now aim to develop and make available to the international community a comprehensive catalog of regulatory genomic features and variation in these regions that relate to human diseases. Our project is divided into a data integration (DI) and an information extraction (IE) part. In the last funding period, we developed the first text corpus annotated with regulatory information. This was used to train text-mining algorithms that can detect regulatory sequence elements in new texts. This resulted in the first text-mining-derived collection of these elements and their putative associations with genes and diseases. In addition, we have developed an entity normalization method based on Deep Neural Networks and large language models. In the second application phase, for DI, we will focus on updating and expanding the number of integrated databases and automating the integration process. In IE, we will shift our focus from entity recognition and normalization to entity relationship extraction. Training the models will require extending the annotation of the corpus to represent relationships between regulatory features and genes, variants, and diseases. We will train state-of-the-art methods for relation extraction on this extended corpus and apply the trained models to disease-specific text collections. For curating the results, we will develop an innovative method for rapid annotation that focuses on user satisfaction and ease of use, an aspect that is still understudied in currently available software tools. For quick and easy access to all regulatory feature data integrated, extracted, and curated in the project, we will develop a user-friendly web interface with intuitive visualizations that will be integrated into the RegulationSpotter website.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Learning Table Similarity Measures
-
批准号:388146305
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2017
-
负责人:Professor Dr. Ulf Leser
-
依托单位:
Web Data Analytics and Scientific Workflows
-
批准号:248359967
-
项目类别:Research Units
-
资助金额:$0.0万
-
财政年份:2013
-
负责人:Professor Dr. Ulf Leser
-
依托单位:
Scalable Information Extraction in Stratosphere
-
批准号:174466407
-
项目类别:Research Units
-
资助金额:$0.0万
-
财政年份:2010
-
负责人:Professor Dr. Ulf Leser
-
依托单位:
海外基金