Data Profiling and Data Cleansing on Stratosphere
Data Profiling and Data Cleansing on Stratosphere
批准号:
248360345
负责人:
Professor Dr. Felix Naumann
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Units
财政年份:
2013
资助国家:
德国
项目状态:
已结题
起止时间:
2012-12-31 至 2016-12-31
中文摘要
项目E遵循两个主要研究目标,即“低延迟清理”和“可扩展数据分析”。数据清理代表了Stratosphere的一个真正的、经过测试的应用领域,但实现低延迟和早期结果可以实现新的应用领域。对于开发人员来说,数据概要分析既是一个具有挑战性的应用领域,也是查询规划、数据分析、数据集成和清理的核心技术。低延迟清理旨在减少典型的数据清理和集成操作的管道阻塞特性。例如,要回答一个需要无重复结果的查询,需要调用一个复杂的重复检测操作符,它可能在发出结果之前执行多个排序步骤和一些高级集群。在大数据场景中,并非所有数据最初都是可用的,或者在用户交互场景中,这种延迟是不可接受的。我们计划研究能够在对结果质量影响最小的情况下处理如此高通量数据的专门清洗操作员。例如,输出预先选择的特别干净的数据可以为准备更复杂的清理步骤“争取”必要的时间。可伸缩数据分析的目标是确定关于非常大的数据集的元数据,例如简单的元组计数或列唯一性,以及中等复杂的信息,例如频繁值模式和数据类型,以及难以确定的高度复杂的信息,包括(有条件的)功能依赖关系和包含依赖关系。第3.1节给出了数据分析任务的全面概述。由于通用工具的复杂性,它们通常只覆盖可能的数据分析任务的子集。专门化方法通常更有效,但只涵盖一种类型的信息,并且通常假设所有数据都驻留在主存中。我们的目标是缓解这两个方面:首先,我们不能假设基于主内存的方法就足够了,而是需要处理必须分布在多个站点上的数据。其次,我们计划将各种任务的计算结合起来,以最小化I/O活动。数据概要分析的结果至少有两个直接用途。首先,它们作为Stratosphere统计组件的输入,该组件反过来支持查询优化。其次,数据分析结果作为数据清理方法的输入。例如,关于列中的公共值模式的知识允许在该列上配置规范化操作符。反之亦然,一些简单的数据清理。
英文摘要
Project E follows two main research goals, namely “Low-Latency Cleansing” and “Scalable Data Profiling”. Data cleansing represents a true and tested application area of Stratosphere, yet achieving lowlatency and early results enables new application areas. Data profiling serves both as a challenging application area for developers and as a core technology for query planning, data analysis, and data integration and cleansing.Low-latency cleansing aims at reducing the pipeline- blocking nature that is typical of data cleansing and integration operators. For instance, to answer a query requiring a duplicate-free result, a complex duplicate detection operator is invoked, which might perform multiple sorting steps and possibly some advanced clustering before emitting results. In big-data scenarios where not all data is even initially available or in user-interactive scenarios such latency is unacceptable. We plan to investigate specialized cleansing operators that can handle such high-throughput data with minimal effect on thequality of the outcome. For instance, outputting a pre-selection of particularly clean data can “buy” the necessary time to prepare more complex cleansing steps.Scalable data profiling has the goal to determine metadata about very large datasets, such as simple tuple-counts or column-uniqueness, but also moderately complex information, such as frequent value patterns and data types, to information that is highly complex to determine, including (conditional) functional dependencies and inclusion dependencies. Section 3.1 gives a comprehensive overview of data profiling tasks. General tools usually cover only a subset of possible data profiling tasks, due to their complexity. Specialized methods are often more efficient, but cover only one type of information and often assume that all data resides in main memory. Our goal is to alleviate both aspects: First, we cannot assume that main memory-based approaches suffice but rather need to handle data that must be distributed over multiple sites. Second, we plan to combine calculations for various tasks so as to minimize I/O activity. The results of data profiling have at least two immediate uses. First, they serve as input to Stratosphere’s statistics component, which in turn supports query optimization. Second, data profiling results serve as input to data cleansing methods. For instance, knowledge about common value patterns in a column allows the configuration of a normalization operator on that column. Vice-versa, some simple data cleansing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Uncertainty and Data Cleansing in the Stratosphere Cloud Data Management System
-
批准号:174473156
-
项目类别:Research Units
-
资助金额:$0.0万
-
财政年份:2010
-
负责人:Professor Dr. Felix Naumann
-
依托单位:
Merging autonomous content for the effective integration of autonomous and heterogeneous information sources
-
批准号:5401631
-
项目类别:Independent Junior Research Groups
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Professor Dr. Felix Naumann
-
依托单位:
海外基金