A journey to Semantic Web query federation in the life sciences.

A journey to Semantic Web query federation in the life sciences.
复制标题

DOI:
10.1186/1471-2105-10-s10-s10
复制
发表时间:
2009-10-01
期刊:
影响因子:
3
通讯作者:
Paschke A
Paschke A
中科院分区:
生物学4区
文献类型:
--
作者:
Cheung KH;Frost HR;Marshall MS;Prud'hommeaux E;Samwald M;Zhao J;Paschke A

文献摘要

被引文献

相似文献

随着人们对在生物医学领域采用语义网的兴趣不断增长,语义网技术也在不断发展和成熟。近年来出现了各种各样的技术方法,包括三重存储技术、SPARQL端点、关联数据和互联数据集词汇表。除了数据仓库构造之外,这些技术方法还可用于支持动态查询联合。作为一个社区的努力,BioRDF任务小组,在医疗保健和生命科学兴趣小组的语义网,正在探索如何利用这些新兴的方法来执行跨不同神经科学数据源的分布式查询。我们创建了两个医疗保健和生命科学知识库。我们探索了各种语义Web方法来描述、映射和动态查询多个数据集。我们已经展示了几种联合方法,这些方法整合了关于神经元和受体的不同类型的信息,这些信息在基础、临床和转化神经科学研究中起着重要作用。特别是,我们创建了一个原型受体资源管理器,它使用OWL映射来提供一个集成的受体列表,并针对不同的SPARQL端点执行单独的查询。我们还使用了AIDA工具包,它是针对知识工作者群体的,这些知识工作者协作地搜索、注释、解释和丰富来自不同位置的大型异构文档集合。我们已经研究了一个名为“联邦”的工具,它允许将全局SPARQL查询分解为针对提供SPARQL或SQL查询接口的远程数据库的子查询。最后,我们探讨了如何使用互连数据集词汇表(voiD)来创建元数据,用于描述作为关联数据uri或SPARQL端点公开的数据集。我们已经演示了如何使用一组新颖的、最先进的语义Web技术来支持神经科学查询联合场景。我们已经确定了这些技术的优点和缺点。虽然语义Web提供了一个包括使用统一资源标识符(Uniform Resource Identifiers, URI)在内的全局数据模型,但语义等效URI的激增阻碍了大规模数据集成。我们的工作有助于指导研究和工具开发,这将有利于这个社区。
As interest in adopting the Semantic Web in the biomedical domain continues to grow, Semantic Web technology has been evolving and maturing. A variety of technological approaches including triplestore technologies, SPARQL endpoints, Linked Data, and Vocabulary of Interlinked Datasets have emerged in recent years. In addition to the data warehouse construction, these technological approaches can be used to support dynamic query federation. As a community effort, the BioRDF task force, within the Semantic Web for Health Care and Life Sciences Interest Group, is exploring how these emerging approaches can be utilized to execute distributed queries across different neuroscience data sources. We have created two health care and life science knowledge bases. We have explored a variety of Semantic Web approaches to describe, map, and dynamically query multiple datasets. We have demonstrated several federation approaches that integrate diverse types of information about neurons and receptors that play an important role in basic, clinical, and translational neuroscience research. Particularly, we have created a prototype receptor explorer which uses OWL mappings to provide an integrated list of receptors and executes individual queries against different SPARQL endpoints. We have also employed the AIDA Toolkit, which is directed at groups of knowledge workers who cooperatively search, annotate, interpret, and enrich large collections of heterogeneous documents from diverse locations. We have explored a tool called "FeDeRate", which enables a global SPARQL query to be decomposed into subqueries against the remote databases offering either SPARQL or SQL query interfaces. Finally, we have explored how to use the vocabulary of interlinked Datasets (voiD) to create metadata for describing datasets exposed as Linked Data URIs or SPARQL endpoints. We have demonstrated the use of a set of novel and state-of-the-art Semantic Web technologies in support of a neuroscience query federation scenario. We have identified both the strengths and weaknesses of these technologies. While Semantic Web offers a global data model including the use of Uniform Resource Identifiers (URI's), the proliferation of semantically-equivalent URI's hinders large scale data integration. Our work helps direct research and tool development, which will be of benefit to this community.