Open Data Integration

Open Data Integration
复制标题

开放数据集成

DOI:
--
复制
发表时间:
2018
影响因子:
2.5
通讯作者:
Renée J. Miller
Renée J. Miller
中科院分区:
计算机科学2区
文献类型:
--
作者:
Renée J. Miller

文献摘要

被引文献

相似文献

开放数据在支持政府和组织透明度方面发挥着重要作用。许多组织正在采用 开放数据原则 承诺使他们的开放数据完整、原始和及时。这些属性使得这些数据对数据科学家来说非常有价值。然而,科学家们通常没有 先验 关于哪些数据可用的知识(其模式或内容)。尽管如此,他们希望能够使用开放数据,并将其与他们正在研究的其他公共或私人数据集成。传统上,数据集成是使用称为 查询发现 其中主要任务是发现将数据从一种形式转换为另一种形式的查询(或转换)。目标是找到正确的操作符来将数据连接、嵌套、分组、链接和扭曲成所需的形式。我们引入了一种新的集成思维模式, 数据发现, 但由数据分析需求驱动的高效互联网规模的发现。我们描述了一个研究议程和最近的进展,在开发可扩展的数据分析或查询感知的数据发现算法,提供高召回率和准确性的大规模数据存储库。
Open data plays a major role in supporting both governmental and organizational transparency. Many organizations are adopting Open Data Principles promising to make their open data complete, primary, and timely. These properties make this data tremendously valuable to data scientists. However, scientists generally do not have a priori knowledge about what data is available (its schema or content). Nevertheless, they want to be able to use open data and integrate it with other public or private data they are studying. Traditionally, data integration is done using a framework called query discovery where the main task is to discover a query (or transformation) that translates data from one form into another. The goal is to find the right operators to join, nest, group, link, and twist data into a desired form. We introduce a new paradigm for thinking about integration where the focus is on data discovery, but highly efficient internet-scale discovery that is driven by data analysis needs. We describe a research agenda and recent progress in developing scalable data-analysis or query-aware data discovery algorithms that provide high recall and accuracy over massive data repositories.