Datenintegration

Datenintegration
复制标题

数据整合

DOI:
--
复制
发表时间:
2012
期刊:
it - Information Technology
影响因子:
--
通讯作者:
Melanie Herschel
Melanie Herschel
中科院分区:
--
文献类型:
--
作者:
Melanie Herschel

文献摘要

被引文献

相似文献

数据集成的目标是将驻留在分布式、自治和异构数据库中的数据组合到一个一致的数据视图中。它的应用非常广泛,从科学数据的数据集成(帮助科学家共享、理解、重用和补充过去的成果)到企业中的数据集成(例如建立数据仓库、执行商业智能或实施主数据管理)以及Web上的数据集成(比较在线购物或链接开放数据)。为了实现数据集成,有三个主要问题一直是数据库研究界和IT行业特别感兴趣的。首先,必须克服数据源的数据模型和模式之间的异构性。第二,数据源可能在真实世界实体的集合中重叠,例如它们所表示的人或产品,并且需要识别它们对同一实体的多个且通常不同的表示(所谓的重复)。最后,在集成的结果中,每个实体都应该被精确地表示一次,因此需要将副本合并到单个表示中,这个问题也被称为数据融合。这期特刊从研究和工业的角度介绍了与数据集成相关的一些解决方案和新挑战。Mecca和Papotti的文章描述了模式映射和数据交换解决方案的最新技术水平,通过桥接数据模型和要集成的数据源模式之间的异构性来解决上述第一个步骤。模式映射技术由于其声明性、清晰的语义、易于使用的设计工具以及在部署步骤中的效率和模块化而获得了极大的普及。本文将调查的方法分为三个时代:产生理论基础和早期工具的英雄时代,模式映射工具已发展到复杂系统中并已转化为商业和开源工具的银时代,以及即将到来的黄金时代,新的研究机会和新一代的系统能够处理更大的一类真实的-生活应用Maier、Oberhofer和施瓦茨描述了一种商业化的数据集成方法,它解决了上述三个主要问题,并已在世界范围内的大型数据集成项目中使用。这种方法基于这样的观察:典型数据集成工作的最大部分致力于在健壮和高性能的商业系统中实现转换、清理和数据验证逻辑。这项工作很简单,不需要超出商业产品知识的技能,但它是非常劳动密集型和容易出错。他们的方法有助于数据集成项目的工业化,并大大降低了简单但劳动密集型的工作量。其关键思想是,数据集成项目的目标环境具有预定义的数据模型和关联的Meta数据,可以利用这些数据来构建和自动化数据集成过程。在她关于多尺度数据整合的文章中,BertiEquille提出了一些具有挑战性的研究方向,用于整合来自观测科学领域的大量多尺度科学数据。这些数据被集中收集,以测量地球的各种属性。例如,科学家观察环境条件,生态系统或生物物种。理解全球变暖等复杂现象和根据时空数据预测趋势的能力已成为观测科学的一个主要问题,而多尺度数据整合方面的理论和技术进步对此至关重要。本文描述了观测科学中数据集成的几个用例,并概述了由于时间,空间,结构,语义或分析依赖性,从原始测量数据到处理数据和衍生统计数据的不同级别的数据粒度或数据抽象,取决于学科的各种数据解释或使用,时空数据的质量异质性以及缩放问题而带来的挑战。
Data integration aims at combining data that resides in distributed, autonomous, and heterogeneous databases into a single consistent view of the data. Its applications are abundant, ranging from data integration in scientific data (helping scientists share, understand, reuse and complement past results) to data integration in enterprises (for instance to set up data warehouses, perform business intelligence or implement master data management) and data integration on the Web (comparison online shopping or linking open data). In order to achieve data integration, three major problems have been of particular interest to both the database research community and IT industry. First, the heterogeneity between data models and schemas of data sources has to be overcome. Second, data sources may overlap in the sets of real-world entities such as persons or products they represent and their multiple and usually different representations of a same entity, so called duplicates, need to be identified. Finally, in the integrated result, every entity should be represented exactly once, so duplicates need to be merged to a single representation, a problem also referred to as data fusion. This special issue covers some solutions and new challenges related to data integration, both from a research and an industrial perspective. The article by Mecca and Papotti describes the state-ofthe art of schema mapping and data exchange solutions, employed to address the first of the above steps by bridging the heterogeneity between data models and schemas of data sources to be integrated. Schema mapping techniques have acquired great popularity due to their declarative nature, clean semantics, easy to use design tools and their efficiency and modularity in the deployment step. The article divides the surveyed approaches into three ages: the heroic age that produced the theoretical foundations and early tools, the silver age when schema mapping tools have grown their way into complex systems and have been translated into both commercial and open-source tools, and a forthcoming golden age with novel research opportunities and a new generation of systems capable of dealing with a significantly larger class of real-life applications. Maier, Oberhofer and Schwarz describe a commercial approach to data integration that addresses the three main problems above (among others) and that has been used in large data integration projects worldwide. This approach is based on the observation that the largest part of a typical data integration effort is dedicated to the implementation of transformation, cleansing, and data validation logic in robust and highly performing commercial systems. This effort is simple and does not demand skills beyond commercial product knowledge, but it is very labour-intensive and error prone. Their approach helps to industrialize data integration projects and significantly lowers the amount of simple, but labourintensive work. The key idea is that the target landscape for a data integration project has pre-defined data models and associated meta data which can be leveraged for building and automating the data integration process. In her article on multi-scale data integration, BertiEquille presents some challenging research directions for integrating massive multi-scale scientific data from the observational science domain. This data is intensively collected in order to measure various properties of the Earth. For instance, scientists observe environmental conditions, ecosystems, or biological species. The ability to understand complex phenomena such as global warming and to predict trends from spatio-temporal data have become a major issue in observational science, for which theoretical and technical advances in multi-scale data integration are essential. The paper describes several use cases of data integration in observational sciences and outlines challenge due to temporal, spatial, structural, semantic or analytic dependencies, different levels of data granularity or data abstraction from raw measurement data to processed data and derived statistics, various data interpretations or usages depending on the disciplines, quality heterogeneity of spatio-temporal data, and scaling issues.