Ontology-Based Querying with Bio2RDF's Linked Open Data.

Ontology-Based Querying with Bio2RDF's Linked Open Data.
复制标题

DOI:
10.1186/2041-1480-4-s1-s1
复制
发表时间:
2013-04-15
影响因子:
1.9
通讯作者:
Dumontier, Michel
Dumontier, Michel
中科院分区:
工程技术4区
文献类型:
--
作者:
Callahan, Alison;Cruz-Toledo, Jose;Dumontier, Michel

文献摘要

被引文献

相似文献

背景技术背景:在这个后“组学”时代,生命科学家的一项关键活动是从众多独立的数据库中搜索和整合生物数据。然而,我们查找相关数据的能力受到了由多种数据格式支持的非标准Web和数据库接口的阻碍。这种异质性提出了一个压倒性的障碍,以发现和重用的资源已开发在巨大的公共expensed.To解决这个问题,开源Bio2RDF项目促进了一个简单的公约,使用语义Web技术集成不同的生物数据。然而,查询Bio2RDF仍然很困难,由于缺乏统一的表示Bio2RDF datasets.RESULTS:我们描述了更新Bio2RDF,包括更紧密的集成在19个新的和更新的RDF数据集。所有可用的开源脚本首先被整合到单个GitHub存储库中,然后使用通用的API进行重新开发,该API使用集中式数据集注册表生成规范化的IRI。然后,我们映射数据集的特定类型和关系的Semanticscience集成本体(SIO),并展示了简化的联邦查询跨多个Bio2RDF endpoints.CONCLUSIONS:这个协调发布标志着Bio2RDF开源链接数据框架的一个重要里程碑。主要是,它提高了Bio2RDF网络中链接数据的质量,并使本地访问或重新创建链接数据变得更容易。我们希望通过确定优先数据库并增加词汇覆盖率以增加SIO之外的其他数据集词汇来继续改进链接数据的Bio2RDF网络。
BACKGROUND: A key activity for life scientists in this post "-omics" age involves searching for and integrating biological data from a multitude of independent databases. However, our ability to find relevant data is hampered by non-standard web and database interfaces backed by an enormous variety of data formats. This heterogeneity presents an overwhelming barrier to the discovery and reuse of resources which have been developed at great public expense.To address this issue, the open-source Bio2RDF project promotes a simple convention to integrate diverse biological data using Semantic Web technologies. However, querying Bio2RDF remains difficult due to the lack of uniformity in the representation of Bio2RDF datasets.RESULTS: We describe an update to Bio2RDF that includes tighter integration across 19 new and updated RDF datasets. All available open-source scripts were first consolidated to a single GitHub repository and then redeveloped using a common API that generates normalized IRIs using a centralized dataset registry. We then mapped dataset specific types and relations to the Semanticscience Integrated Ontology (SIO) and demonstrate simplified federated queries across multiple Bio2RDF endpoints.CONCLUSIONS: This coordinated release marks an important milestone for the Bio2RDF open source linked data framework. Principally, it improves the quality of linked data in the Bio2RDF network and makes it easier to access or recreate the linked data locally. We hope to continue improving the Bio2RDF network of linked data by identifying priority databases and increasing the vocabulary coverage to additional dataset vocabularies beyond SIO.