课题基金 / 基金详情

Information Integration of Heterogeneous Data Sources

Information Integration of Heterogeneous Data Sources
异构数据源信息集成
批准号:
7481818
负责人:
Mansur R. Kabuka
金额:
$47.34万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-03-15 至 2011-03-31

项目摘要

项目成果

Mansur R. Kabuka的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):不断产生的丰富的生物和生物医学数据保证了生命科学的巨大进步。为了实现这一前景,这一快速增长的信息库需要被有效地整合,也就是说,以这样一种方式结合起来,即可以对其进行查询以提取相关数据,这些数据随后可以被分析以回答有意义的研究问题。这项提议的主要目标是开发GeneTegra系统,这是一个信息集成解决方案,提供一个共同的交互环境,以查询来自多个来源的数据和知识。为了有效整合来自不同数据源的知识,必须克服两个主要障碍:句法异质性,其中数据源具有不同的表示和访问机制;以及语义可变性,其中相似的词汇术语可能涉及多个概念,而不同的术语可能涉及同一概念。GeneTegra系统通过使用语义Web技术来解决这些障碍:使用Web Ontology Language(OWL)作为不同格式的数据源的公共数据和知识表示构建的本体,用于生成和维护这些本体表示的自动化机制,以及基于可重用的、面向服务的中介的健壮的系统架构。该系统的核心包括在本项目第一阶段开发的通用算法、过程和机制,这些算法、过程和机制能够自动生成本体,自动识别本体模型之间的语义对应关系,以及在这些本体建模的、分布式的、异构源上创建和执行查询。在第二阶段,GeneTegra系统将作为一个以人为中心的解决方案进行开发、实施和测试,该解决方案建立在第一阶段开发的核心组件之上,包括一个用于查询创建和执行的高度可用的界面、一个使用网络服务标准注册、共享和重新使用信息的机制、一个用于确定数据质量和查询可靠性的机制、以及一个安全和隐私子系统,该子系统允许建立协作社区,同时确保用户通过该系统获得适当的身份验证和授权访问信息。GeneTegra系统的设计和评估将具体解决与调查基因-表型相关的来源以及识别导致人类疾病和状况的基因相关的来源的整合问题。公共卫生相关性GeneTegra系统是一个信息集成解决方案,它提供了一个通用的交互环境来查询来自多个不同来源的数据和知识。它使用本体作为语义和语法建模的基本公式,并包含用于生成这些本体以及用于重用和共享集成配置的自动化机制。它专门用于解决与调查基因-表型相关的来源的综合查询,以及与识别导致人类疾病和状况的基因有关的问题。
英文摘要
DESCRIPTION (provided by applicant): The wealth of biological and biomedical data constantly being generated promises dramatic advancement in the life sciences. To realize this promise, this pool of rapidly expanding information needs to be efficiently integrated, that is, combined in such a way that it can be queried to extract relevant data that can be subsequently analyzed to answer meaningful research questions. The main objective of this proposal is to develop the GeneTegra System, an information integration solution that provides a common interaction environment to query data and knowledge from multiple sources. Two main obstacles have to be overcome in order to attain an effective integration of knowledge from different data sources: syntactic heterogeneity, where data sources have different representation and access mechanisms; and semantic variability, where similar lexical terms may refer to multiple concepts or dissimilar terms refer to the same concept. The GeneTegra System addresses these obstacles through the use of Semantic Web technologies: ontologies constructed using the Web Ontology Language (OWL) as a common data and knowledge representation for data sources of diverse formats, automated mechanisms for the generation and maintenance of these ontology representations, and a robust system architecture based on reusable, service-oriented mediators. The core of the proposed system consists of general algorithms, procedures, and mechanisms developed during Phase I of this project, that enable the automatic generation of ontologies, the automated identification of semantic correspondences between ontology models, and the creation and execution of queries over these ontology- modeled, distributed, heterogeneous sources. In Phase II, the GeneTegra System will be developed, implemented, and tested as a human-centered solution building on the core components developed during Phase I, incorporating a highly usable interface for query creation and execution, a mechanism for registration, sharing, and re-use of information using Web Services standards, a mechanism for determining quality of data and query reliability, and a security and privacy subsystem that allows the construction of collaborative communities while ensuring that users are properly authenticated and authorized to access information through the system. The GeneTegra System will be designed and evaluated to specifically address the integration of sources relevant to investigations of genotype-phenotype associations and to the identification of genes responsible for human diseases and conditions. PUBLIC HEALTH RELEVANCE The GeneTegra System is an information integration solution that provides a common interaction environment to query data and knowledge from multiple heterogeneous sources. It uses ontologies as the base formulism for semantic and syntactic modeling, and contains automated mechanisms for the generation of these ontologies, and for the reuse and sharing of integration configurations. It is specifically designed to address the integrated querying of sources relevant to investigations of genotype-phenotype associations and to the identification of genes responsible for human diseases and conditions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Deep-CDS: Deep Learning Semantic Data Lake for Clinical Decision Support
  • 批准号:
    10546333
  • 项目类别:
  • 资助金额:
    $25.35万
  • 财政年份:
    2022
  • 负责人:
    Mansur R. Kabuka
  • 依托单位:
Deep-CDS: Deep Learning Semantic Data Lake for Clinical Decision Support
  • 批准号:
    10747223
  • 项目类别:
  • 资助金额:
    $71.43万
  • 财政年份:
    2022
  • 负责人:
    Mansur R. Kabuka
  • 依托单位:
Microbiome Meta-Analysis Platform
  • 批准号:
    10011865
  • 项目类别:
  • 资助金额:
    $29.38万
  • 财政年份:
    2017
  • 负责人:
    Mansur R. Kabuka
  • 依托单位:
Semantic Data Lake for Biomedical Research
  • 批准号:
    9536289
  • 项目类别:
  • 资助金额:
    $58.97万
  • 财政年份:
    2016
  • 负责人:
    Mansur R. Kabuka
  • 依托单位:
海外基金