课题基金 / 基金详情

Big Data Cleaning

Big Data Cleaning
大数据清洗
批准号:
RGPIN-2015-06552
负责人:
Szlichta, Jaroslaw
金额:
$1.31万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31
关键词:

项目摘要

项目成果

Szlichta, Jaroslaw的其他基金

相似基金

相关文献

中文摘要
翻译
由于数据质量差,组织越来越难以从数据中获得价值。数据质量差是基于数据进行有效和高质量决策的障碍。完整性约束(业务规则)是用于保持数据完整性的基本工具。它们指定了应该在数据上保持的域语义。声明式数据清理已成为评估和提高数据质量的有效工具。我们将解决一些重要的挑战,在将声明式数据清洗应用于大数据时,由于这些数据的规模,复杂性和巨大的异质性而出现。鉴于现代应用程序中存在的不同数据格式和数据表示的激增,这些异构数据源的集成会导致现有方法无法处理的微妙不一致性。首先,考虑到大数据的动态特性,我们将为动态数据环境开发新的连续数据清理方法。随着数据和约束的发展,我们需要根据这些增量更改来确定修复,而不必每次都从头开始修复过程。其次,我们将研究度量和本体约束,以增强声明性数据清理,并探索具有可扩展规则规范的整体数据清理。我们将开发一个新的框架,结合统计(采用统计学原理)和逻辑(使用各种形式的逻辑推理声明依赖)数据清理。最近在数据清理方面的工作提出了很大程度上只与这些领域之一隔离的解决方案。第三,认识到大数据的巨大异质性以及自动化很少能提供100%的准确性,我们将开发新技术来解释数据清理解决方案的起源(血统)。起源帮助用户理解为什么(以及如何)得出清理决策,并使用户能够调试和纠正自动化解决方案。我们还将使用数据出处来指导清理算法选择更准确的修复。*拟议的研究是有益的加拿大(和国际)的商业和政府组织。组织可以通过对高质量数据进行分析决策来显著受益。我们的解决方案可用于医疗保健、电信(例如,罗杰斯和AT&T),金融(例如,蒙特利尔银行和TD银行)和政府(例如,加拿大统计局和交通部)机构。这项研究的结果也将是感兴趣的软件供应商,如IBM,甲骨文,SAP和微软。该计划将培训学生数据管理系统,并使他们在申请学术界和工业界的工作时处于竞争地位。我们预计多达10名学生(包括本科生)将在本研究计划中接受培训。
英文摘要
Organizations are finding it increasingly difficult to reap value from their data due to poor data quality. Poor data quality is a barrier to effective and high quality decision-making based on data. Integrity constraints (business rules) are the fundamental tools used to preserve data integrity. They specify the domain semantics that should hold over the data. Declarative data cleaning has emerged as an effective tool for both assessing and improving the quality of data.****We will address some important challenges, in applying declarative data cleaning to big data, that arise due to the scale, complexity, and massive heterogeneity of such data. Given the proliferation of different data formats and data representations that exist in modern applications, the integration of these heterogeneous data sources leads to subtle inconsistencies that are not handled by the existing methods.****First, given the dynamic nature of big data, we will develop new continuous data cleaning methods for dynamic data environments. As the data and constraints evolve we need to identify repairs based on this incremental changes, without having to start the repair process from scratch each time. Second, we will investigate metrical and ontological constrains to enhance declarative data cleaning and explore holistic data cleaning with an extensible rule specification. We will develop a novel framework that combines statistical (employing principles of statistics) and logical (using various forms of logical reasoning over declarative dependencies) data cleaning. Recent work in data cleaning has proposed solutions largely isolated to only one of these areas. Third, recognizing the massive heterogeneity of big data and that automation will rarely provide 100% accuracy, we will develop new techniques to explain the provenance (lineage) of data cleaning solutions. Provenance helps users to understand why (and how) a cleaning decision was derived and can enable users to debug and correct automated solutions. We will also use data provenance to guide a cleaning algorithm to choose more accurate repairs.***The proposed research is beneficial for Canadian (and international) business and government organizations. Organizations can significantly benefit by making analytical decisions over high quality data. Our solutions can be utilized in healthcare, telecommunication (e.g., Rogers and AT&T), financial (e.g., the Bank of Montreal and TD Bank) and governmental (e.g., Statistics Canada and the Ministry of Transportation) institutions. The outcomes of this research will also be of interest to software vendors, such as IBM, Oracle, SAP and Microsoft. The proposed program will train students in data management systems and place them in a competitive position while applying for jobs in academia and industry. We expect up to ten students (including undergraduate students) to receive training in this research program.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.42万
  • 财政年份:
    2022
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $0.13万
  • 财政年份:
    2022
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.55万
  • 财政年份:
    2021
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.55万
  • 财政年份:
    2020
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
  • 批准号:
    61373035
  • 项目类别:
    面上项目
  • 资助金额:
    77.0万元
  • 批准年份:
    2013
  • 负责人:
    冯志勇
  • 依托单位: