课题基金 / 基金详情

Big Data Cleaning

Big Data Cleaning
大数据清洗
批准号:
RGPIN-2015-06552
负责人:
Szlichta, Jaroslaw
金额:
$1.31万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31
关键词:

项目摘要

项目成果

Szlichta, Jaroslaw的其他基金

相似基金

相关文献

中文摘要
翻译
组织发现,由于数据质量差,从其数据中获取价值变得越来越困难。数据质量差是基于数据进行有效和高质量决策的障碍。完整性约束(业务规则)是用于维护数据完整性的基本工具。它们指定了应该控制数据的域语义。声明性数据清理已经成为评估和提高数据质量的有效工具。*我们将解决由于大数据的规模、复杂性和海量异构性而产生的声明性数据清理在大数据中应用的一些重要挑战。考虑到现代应用中存在的不同数据格式和数据表示的激增,这些异构数据源的集成导致了现有方法无法处理的微妙不一致。*首先,鉴于大数据的动态性质,我们将为动态数据环境开发新的持续数据清理方法。随着数据和约束的发展,我们需要基于这种增量更改来确定修复,而不必每次都从头开始修复过程。其次,我们将研究度量约束和本体约束,以增强声明性数据清理,并探索具有可扩展规则规范的整体数据清理。我们将开发一个新的框架,将统计(使用统计原理)和逻辑(使用各种形式的逻辑推理而不是声明性依赖)相结合,进行数据清理。最近在数据清理方面的工作提出了在很大程度上孤立于其中一个领域的解决方案。第三,认识到大数据的海量异构性和自动化很少提供100%的准确性,我们将开发新的技术来解释数据清理解决方案的起源(谱系)。Provenance帮助用户了解为什么(以及如何)得出清洁决策,并使用户能够调试和纠正自动化解决方案。我们还将使用数据来源来指导清洁算法选择更准确的修复。*建议的研究对加拿大(以及国际)企业和政府组织有利。组织可以通过对高质量数据做出分析决策而显著受益。我们的解决方案可用于医疗保健、电信(如罗杰斯和AT&T)、金融机构(如蒙特利尔银行和TD银行)和政府机构(如加拿大统计局和交通部)。这项研究的结果也将引起IBM、甲骨文、SAP和微软等软件供应商的兴趣。拟议中的计划将对学生进行数据管理系统方面的培训,使他们在申请学术界和工业界的工作时处于竞争地位。我们预计最多有10名学生(包括本科生)将接受该研究计划的培训。**
英文摘要
Organizations are finding it increasingly difficult to reap value from their data due to poor data quality. Poor data quality is a barrier to effective and high quality decision-making based on data. Integrity constraints (business rules) are the fundamental tools used to preserve data integrity. They specify the domain semantics that should hold over the data. Declarative data cleaning has emerged as an effective tool for both assessing and improving the quality of data.****We will address some important challenges, in applying declarative data cleaning to big data, that arise due to the scale, complexity, and massive heterogeneity of such data. Given the proliferation of different data formats and data representations that exist in modern applications, the integration of these heterogeneous data sources leads to subtle inconsistencies that are not handled by the existing methods.****First, given the dynamic nature of big data, we will develop new continuous data cleaning methods for dynamic data environments. As the data and constraints evolve we need to identify repairs based on this incremental changes, without having to start the repair process from scratch each time. Second, we will investigate metrical and ontological constrains to enhance declarative data cleaning and explore holistic data cleaning with an extensible rule specification. We will develop a novel framework that combines statistical (employing principles of statistics) and logical (using various forms of logical reasoning over declarative dependencies) data cleaning. Recent work in data cleaning has proposed solutions largely isolated to only one of these areas. Third, recognizing the massive heterogeneity of big data and that automation will rarely provide 100% accuracy, we will develop new techniques to explain the provenance (lineage) of data cleaning solutions. Provenance helps users to understand why (and how) a cleaning decision was derived and can enable users to debug and correct automated solutions. We will also use data provenance to guide a cleaning algorithm to choose more accurate repairs.***The proposed research is beneficial for Canadian (and international) business and government organizations. Organizations can significantly benefit by making analytical decisions over high quality data. Our solutions can be utilized in healthcare, telecommunication (e.g., Rogers and AT&T), financial (e.g., the Bank of Montreal and TD Bank) and governmental (e.g., Statistics Canada and the Ministry of Transportation) institutions. The outcomes of this research will also be of interest to software vendors, such as IBM, Oracle, SAP and Microsoft. The proposed program will train students in data management systems and place them in a competitive position while applying for jobs in academia and industry. We expect up to ten students (including undergraduate students) to receive training in this research program.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.42万
  • 财政年份:
    2022
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $0.13万
  • 财政年份:
    2022
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.55万
  • 财政年份:
    2021
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.55万
  • 财政年份:
    2020
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
  • 批准号:
    61373035
  • 项目类别:
    面上项目
  • 资助金额:
    77.0万元
  • 批准年份:
    2013
  • 负责人:
    冯志勇
  • 依托单位: