课题基金 / 基金详情

Continuous Data Curation

Continuous Data Curation
持续数据管理
批准号:
RGPIN-2020-05160
负责人:
Szlichta, Jaroslaw
金额:
$2.42万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
关键词:

项目摘要

项目成果

Szlichta, Jaroslaw的其他基金

相似基金

相关文献

中文摘要
翻译
随着人们对数据分析的兴趣达到前所未有的高度,数据管理已成为一项严峻的挑战。数据管理包括分析、清理和管理数据,为数据分析做准备。随着大数据的崛起,许多信息技术领导者正在忽视进入数据准备的大数据世界的代价。如果没有经过精选的数据,大数据计划可能会花费更长时间、成本更高,带来的好处也更少。根据经济学人信息部2017年对高管进行的一项调查,存储数据的能力不再是问题,只有不到20%的人表示数据存储存在问题,然而,超过50%的人认为其他数据管理任务存在问题,如对账、整合和清理。福布斯在2017年评估称,数据整理约占数据科学家工作的80%。由于缺乏支持原则性数据准备的工具、科学框架和理论基础,数据管理是非常有问题和耗时的。然而,如果没有有原则的数据管理和准备,新的数据分析洞察力就不能得到信任。拟议的研究将为大规模数据管理开发新的方法和软件工具,重点关注大数据的四个V:容量、速度和多样性和准确性带来的技术挑战。首先,鉴于大数据的动态性,我们将开发新的连续数据剖析方法。人们可以仅仅通过查看一个小数据集来描述它,然而,大数据显然需要自动化(和增量)技术。虽然可用的和潜在有用的数据量持续增长,但人类的认知处理能力是固定的。其次,认识到大数据的海量异构性和自动化很少能提供100%的准确性,我们将调查来源和领域本体对流数据的使用,以增强基于规则的数据清理。第三,我们将通过分布式计算和机器学习来构建高效的问题确定和自适应数据库管理系统调优工具。这些新的数据管理技术对政府和企业组织是有益的。组织可以通过对高质量数据做出分析决策而显著受益。我们的解决方案可用于医疗保健(多伦多综合医院)、电信(AT&T和罗杰斯)、社交媒体行业(Twitter和GitHub)和政府(加拿大统计局)机构。预期的可交付成果也将与IBM、甲骨文和SAP等世界知名数据库公司直接相关。数据管理方面的进展将加强加拿大作为信息和通信技术(信通技术)领导者之一的地位。此外,拟议的研究将创造一个独特的培训环境,学生将在其中获得数据科学方面的宝贵经验,这是全球信通技术领域增长最快的学科之一。
英文摘要
With interest in analysis of data at an all-time high data curation has become a critical challenge. Data curation includes profiling, cleaning and managing data in preparation for data analysis. With the ascendance of big data, many information technology leaders are neglecting the price of admission to the big data world of data preparation. Big data initiatives are likely to take longer, cost more, and deliver fewer benefits without curated data. The ability to store data is not a problem anymore, according to a survey of senior executives conducted by the Economist Intelligence Unit in 2017 less than 20% indicated data storage as a problem, however, more than 50% rated other data management tasks, such as reconciliation, integration and cleaning as problematic. Forbes in 2017 assessed that data curation accounts for around 80% of the work of data scientists. Data curation is so problematic and time consuming because of the lack of tools, scientific frameworks, and theoretical foundations to support principled data preparation. However, without principled data management and preparation, new data analytic insights cannot be trusted. The proposed research will develop novel methods and software tools for large-scale data curation, focusing on the technical challenges arising from the four Vs of big data: Volume, Velocity and Variety and Veracity. First, given the dynamic nature of big data, we will develop new continuous data profiling methods. One could profile a small dataset just by looking at it, however, automated (and incremental) techniques are clearly needed for big data. While the amount of available and potentially useful data keeps growing, human cognitive processing capacity is fixed. Second, recognizing the massive heterogeneity of big data and that automation will rarely provide 100% accuracy, we will investigate the use of provenance and domains ontologies over streaming data to enhance rule-based data cleaning. Third, we will build efficient problem determination and adaptive database management system tuning tools through distributed computing and machine learning. These new data curation techniques are beneficial to governmental and business organizations. Organizations can significantly benefit by making analytical decisions over high quality data. Our solutions can be utilized in healthcare (Toronto General Hospital), telecommunication (AT&T and Rogers), social media industry (Twitter and GitHub) and governmental (Statistics Canada) institutions. The anticipated deliverables will also be of direct interest to world-renowned database companies, such as IBM, Oracle and SAP. Advances in data curation will enhance Canada's position as one of the leaders in Information and Communication Technologies (ICT). Furthermore, the proposed research will create a unique training environment, in which students will acquire sought-after experience in data science, one of the fastest-growing disciplines within ICT worldwide.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $0.13万
  • 财政年份:
    2022
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.55万
  • 财政年份:
    2021
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Continuous Data Curation
  • 批准号:
    RGPIN-2020-05160
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.55万
  • 财政年份:
    2020
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
Big Data Cleaning
  • 批准号:
    RGPIN-2015-06552
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2019
  • 负责人:
    Szlichta, Jaroslaw
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
  • 批准号:
    61373035
  • 项目类别:
    面上项目
  • 资助金额:
    77.0万元
  • 批准年份:
    2013
  • 负责人:
    冯志勇
  • 依托单位: