课题基金 / 基金详情

Entity augmentation and data cleaning for machine learning

Entity augmentation and data cleaning for machine learning
用于机器学习的实体增强和数据清理
批准号:
508081-2016
负责人:
Wang, Jiannan
金额:
$4.37万
依托单位:
依托单位国家:
加拿大
项目类别:
Collaborative Research and Development Grants
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Wang, Jiannan的其他基金

相似基金

相关文献

中文摘要
翻译
随着大数据的兴起,公司和组织越来越渴望使用机器学习从数据中提取价值,并实现数据驱动的决策。然而,机器学习通常假设数据已经准备好了,并将其主要重点放在学习和基于数据进行预测上。但是,在现实中,数据通常来自多个来源,并且在数据集成上花费了大量时间;真实世界的数据通常是脏的,数据清理是一个非常耗时和昂贵的过程。根据对数据科学家的采访,他们可以将80%的时间花在数据准备上。在新兴的大数据场景中,当数据量不断增加,或者当数据来自各种各样的来源时,这个问题将进一步加剧。**为此,在这个项目中,我们研究如何降低机器学习的数据准备成本。我们将特别关注两个具有挑战性的研究主题:(1)“实体增强”研究如何有效地从外部数据源增强具有新属性(例如,位置,职业)的实体(例如,餐馆,人员)。(2)“机器学习的数据清洗”研究如何通过只清洗最有利于预测的数据来降低成本。这个项目对加拿大经济有多方面的好处。首先,加拿大越来越多的公司依靠机器学习来做出关键的商业决策(例如,客户流失预测、欺诈检测)。在这个项目中开发的技术可以节省他们的时间来更好地准备用于机器学习的数据,帮助他们提高预测准确性并增加收入。其次,该项目的成果将进一步推动数据科学技术的发展,使小型公司的机器学习民主化,并有助于在加拿大创造更多与数据科学相关的工作岗位
英文摘要
With the rise of Big Data, companies and organizations are increasingly eager to use machine learning to extract value from their data and to enable data-driven decision making. However, machine learning often assumes that data has been well-prepared, and puts its main focus on learning and making predictions based on the data. But, in reality, data often comes from multiple sources and a lot of time is spent on data integration; real-world data is often dirty and data cleaning is an extremely time-consuming and expensive process. According to the interviews of data scientists, they can spend 80% of their time on data preparation. This problem will be further exacerbated in emerging Big Data scenarios when data volumes are increasing, or when data comes from a larger variety of sources.**To this end, in this project, we study how to reduce the cost of data preparation for machine learning. We will particularly focus on two challenging research topics: (1) "Entity augmentation" studies how to efficiently augment entities (e.g., restaurants, persons) with new attributes (e.g., location, occupation) from external data sources. (2) "Data cleaning for machine learning" studies how to reduce the cost by only cleaning the data that are most beneficial to predictions. This project has benefits to the Canadian economy in multiple aspects. First, more and more companies in Canada are relying on machine learning to make critical business decisions (e.g., churn prediction, fraud detection). The techniques developed in this project can save their time to better prepare data for use in machine learning, helping them to improve prediction accuracy and grow revenue. Second, the outcome of the project will further boost the development of data science technologies, democratize machine learning for small companies, and help to create more data science related jobs in Canada.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
DataPrep: Human-in-the-Loop Data Preparation
  • 批准号:
    RGPIN-2021-03995
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.5万
  • 财政年份:
    2022
  • 负责人:
    Wang, Jiannan
  • 依托单位:
DataPrep: Human-in-the-Loop Data Preparation
  • 批准号:
    RGPIN-2021-03995
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.5万
  • 财政年份:
    2021
  • 负责人:
    Wang, Jiannan
  • 依托单位:
Crowdsourced Data Cleaning
  • 批准号:
    RGPIN-2016-05555
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.62万
  • 财政年份:
    2020
  • 负责人:
    Wang, Jiannan
  • 依托单位:
Entity augmentation and data cleaning for machine learning
  • 批准号:
    508081-2016
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $4.37万
  • 财政年份:
    2019
  • 负责人:
    Wang, Jiannan
  • 依托单位:
海外基金