Techniques for optimization of queries on integrated biological resources.

Techniques for optimization of queries on integrated biological resources.
复制标题

综合生物资源查询优化技术。

DOI:
10.1142/s0219720004000648
复制
发表时间:
2004
影响因子:
1
通讯作者:
Eckman,BarbaraA
Eckman,BarbaraA
中科院分区:
生物学4区
文献类型:
--
作者:
Lacroix,Zoé;Raschid,Louiqa;Eckman,BarbaraA

文献摘要

相似文献

今天,科学数据不可避免地被数字化,以各种各样的格式存储,并可通过互联网访问。科学发现越来越多地涉及访问多个异构数据源,集成复杂查询的结果,并应用进一步的分析和可视化应用程序,以收集感兴趣的数据集。构建科学集成平台以支持这些关键任务需要访问和操作从平面文件或数据库中提取的数据、从Web检索的文档以及在仓库中本地物化或由软件生成的数据。现有办法缺乏效率,可能会严重影响这一进程,导致在获取关键资源时出现长时间拖延,或系统无法报告任何结果。有些查询需要很长时间才能得到回答,结果是通过电子邮件返回的,这使得它们与其他结果的集成成为一项繁琐的任务。本文提出了几个需要解决的问题,以提供无缝和有效的整合生物分子数据。已确定的挑战包括:捕获和表示由包括序列或文本搜索引擎和传统查询处理的源支持的各种域特定计算能力;开发用于获取和表示关于源内容、源内容中的重叠和访问成本的语义知识和元数据的方法;开发基于成本和语义的决策支持工具以选择源和能力,并生成有效的查询评估计划。
Today, scientific data are inevitably digitized, stored in a wide variety of formats, and are accessible over the Internet. Scientific discovery increasingly involves accessing multiple heterogeneous data sources, integrating the results of complex queries, and applying further analysis and visualization applications in order to collect datasets of interest. Building a scientific integration platform to support these critical tasks requires accessing and manipulating data extracted from flat files or databases, documents retrieved from the Web, as well as data that are locally materialized in warehouses or generated by software. The lack of efficiency of existing approaches can significantly affect the process with lengthy delays while accessing critical resources or with the failure of the system to report any results. Some queries take so much time to be answered that their results are returned via email, making their integration with other results a tedious task. This paper presents several issues that need to be addressed to provide seamless and efficient integration of biomolecular data. Identified challenges include: capturing and representing various domain specific computational capabilities supported by a source including sequence or text search engines and traditional query processing; developing a methodology to acquire and represent semantic knowledge and metadata about source contents, overlap in source contents, and access costs; developing cost and semantics based decision support tools to select sources and capabilities, and to generate efficient query evaluation plans.