课题基金 / 基金详情

Continuous Data Curation

Continuous Data Curation
持续数据管理
批准号:
RGPIN-2020-05160
负责人:
Szlichta, Jaroslaw
金额:
$2.42万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
关键词:

项目摘要

项目成果

Szlichta, Jaroslaw的其他基金

相似基金

相关文献

中文摘要
翻译
随着人们对数据分析的兴趣空前高涨,数据管理已经成为一个关键的挑战。数据管理包括分析、清理和管理数据,为数据分析做准备。随着大数据的崛起,许多信息技术领导者忽视了数据准备进入大数据世界的代价。如果没有精心策划的数据,大数据计划可能需要更长的时间,成本更高,带来的好处也更少。存储数据的能力不再是一个问题,根据经济学人智库在2017年对高管进行的一项调查,只有不到20%的高管认为数据存储是一个问题,然而,超过50%的高管认为其他数据管理任务,如对账、整合和清理是一个问题。《福布斯》在2017年评估称,数据管理约占数据科学家工作的80%。由于缺乏工具、科学框架和理论基础来支持有原则的数据准备,数据管理是如此成问题和耗时。然而,如果没有有原则的数据管理和准备,新的数据分析见解是不可信的。拟议的研究将为大规模数据管理开发新的方法和软件工具,重点关注大数据的四个v:体积,速度和多样性以及准确性。首先,考虑到大数据的动态特性,我们将开发新的连续数据分析方法。一个人可以通过观察来分析一个小的数据集,然而,大数据显然需要自动化(和增量)技术。虽然可用和潜在有用数据的数量不断增长,但人类的认知处理能力是固定的。其次,认识到大数据的巨大异质性以及自动化很少能提供100%的准确性,我们将研究在流数据上使用来源和领域本体来增强基于规则的数据清理。第三,我们将通过分布式计算和机器学习构建高效的问题确定和自适应数据库管理系统调优工具。这些新的数据管理技术对政府和商业组织都是有益的。组织可以通过对高质量数据进行分析决策而显著受益。我们的解决方案可用于医疗保健(多伦多综合医院),电信(AT&T和Rogers),社交媒体行业(Twitter和GitHub)和政府(加拿大统计局)机构。预计的成果也将直接影响到世界知名的数据库公司,如IBM、Oracle和SAP。数据管理方面的进步将增强加拿大作为信息和通信技术(ICT)领导者之一的地位。此外,拟议的研究将创造一个独特的培训环境,学生将在数据科学方面获得受欢迎的经验,数据科学是全球信息通信技术中发展最快的学科之一。
英文摘要
With interest in analysis of data at an all-time high data curation has become a critical challenge. Data curation includes profiling, cleaning and managing data in preparation for data analysis. With the ascendance of big data, many information technology leaders are neglecting the price of admission to the big data world of data preparation. Big data initiatives are likely to take longer, cost more, and deliver fewer benefits without curated data. The ability to store data is not a problem anymore, according to a survey of senior executives conducted by the Economist Intelligence Unit in 2017 less than 20% indicated data storage as a problem, however, more than 50% rated other data management tasks, such as reconciliation, integration and cleaning as problematic. Forbes in 2017 assessed that data curation accounts for around 80% of the work of data scientists. Data curation is so problematic and time consuming because of the lack of tools, scientific frameworks, and theoretical foundations to support principled data preparation. However, without principled data management and preparation, new data analytic insights cannot be trusted. The proposed research will develop novel methods and software tools for large-scale data curation, focusing on the technical challenges arising from the four Vs of big data: Volume, Velocity and Variety and Veracity. First, given the dynamic nature of big data, we will develop new continuous data profiling methods. One could profile a small dataset just by looking at it, however, automated (and incremental) techniques are clearly needed for big data. While the amount of available and potentially useful data keeps growing, human cognitive processing capacity is fixed. Second, recognizing the massive heterogeneity of big data and that automation will rarely provide 100% accuracy, we will investigate the use of provenance and domains ontologies over streaming data to enhance rule-based data cleaning. Third, we will build efficient problem determination and adaptive database management system tuning tools through distributed computing and machine learning. These new data curation techniques are beneficial to governmental and business organizations. Organizations can significantly benefit by making analytical decisions over high quality data. Our solutions can be utilized in healthcare (Toronto General Hospital), telecommunication (AT&T and Rogers), social media industry (Twitter and GitHub) and governmental (Statistics Canada) institutions. The anticipated deliverables will also be of direct interest to world-renowned database companies, such as IBM, Oracle and SAP. Advances in data curation will enhance Canada's position as one of the leaders in Information and Communication Technologies (ICT). Furthermore, the proposed research will create a unique training environment, in which students will acquire sought-after experience in data science, one of the fastest-growing disciplines within ICT worldwide.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $0.13万
  • 财政年份:
    2022
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.55万
  • 财政年份:
    2021
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.55万
  • 财政年份:
    2020
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Big Data Cleaning
  • 批准号:
    RGPIN-2015-06552
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2019
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
  • 批准号:
    61373035
  • 项目类别:
    面上项目
  • 资助金额:
    77.0万元
  • 批准年份:
    2013
  • 负责人:
    冯志勇
  • 依托单位: