课题基金 / 基金详情

Entity augmentation and data cleaning for machine learning

Entity augmentation and data cleaning for machine learning
用于机器学习的实体增强和数据清理
批准号:
508081-2016
负责人:
Wang, Jiannan
金额:
$4.37万
依托单位:
依托单位国家:
加拿大
项目类别:
Collaborative Research and Development Grants
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Wang, Jiannan的其他基金

相似基金

相关文献

中文摘要
翻译
随着大数据的兴起,公司和组织越来越渴望使用机器学习来从其数据中提取价值,并支持数据驱动的决策。然而,机器学习通常假设数据已经做好了充分的准备,并将主要重点放在学习和基于数据进行预测上。但是,在现实中,数据往往来自多个来源,在数据集成上花费了大量时间;现实世界中的数据往往是肮脏的,数据清理是一个极其耗时和昂贵的过程。根据数据科学家的采访,他们可以将80%的时间花在数据准备上。在新兴的大数据场景中,当数据量不断增加,或者数据来自更多种类的来源时,这个问题会进一步加剧。**为此,在本项目中,我们研究如何降低机器学习的数据准备成本。我们将特别关注两个具有挑战性的研究主题:(1)实体扩充研究如何从外部数据源中有效地扩充具有新属性(例如位置、职业)的实体(例如餐馆、人员)。(2)“机器学习的数据清洗”研究如何通过只清洗对预测最有利的数据来降低成本。该项目对加拿大经济有多方面的好处。首先,加拿大越来越多的公司依赖机器学习来做出关键的商业决策(例如,流失预测、欺诈检测)。该项目开发的技术可以节省他们的时间,以便更好地准备用于机器学习的数据,帮助他们提高预测精度和增加收入。其次,该项目的成果将进一步推动数据科学技术的发展,使小公司的机器学习民主化,并有助于在加拿大创造更多与数据科学相关的就业机会。**
英文摘要
With the rise of Big Data, companies and organizations are increasingly eager to use machine learning to extract value from their data and to enable data-driven decision making. However, machine learning often assumes that data has been well-prepared, and puts its main focus on learning and making predictions based on the data. But, in reality, data often comes from multiple sources and a lot of time is spent on data integration; real-world data is often dirty and data cleaning is an extremely time-consuming and expensive process. According to the interviews of data scientists, they can spend 80% of their time on data preparation. This problem will be further exacerbated in emerging Big Data scenarios when data volumes are increasing, or when data comes from a larger variety of sources.**To this end, in this project, we study how to reduce the cost of data preparation for machine learning. We will particularly focus on two challenging research topics: (1) "Entity augmentation" studies how to efficiently augment entities (e.g., restaurants, persons) with new attributes (e.g., location, occupation) from external data sources. (2) "Data cleaning for machine learning" studies how to reduce the cost by only cleaning the data that are most beneficial to predictions. This project has benefits to the Canadian economy in multiple aspects. First, more and more companies in Canada are relying on machine learning to make critical business decisions (e.g., churn prediction, fraud detection). The techniques developed in this project can save their time to better prepare data for use in machine learning, helping them to improve prediction accuracy and grow revenue. Second, the outcome of the project will further boost the development of data science technologies, democratize machine learning for small companies, and help to create more data science related jobs in Canada.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
DataPrep: Human-in-the-Loop Data Preparation
  • 批准号:
    RGPIN-2021-03995
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.5万
  • 财政年份:
    2022
  • 负责人:
    Wang, Jiannan
  • 依托单位:
DataPrep: Human-in-the-Loop Data Preparation
  • 批准号:
    RGPIN-2021-03995
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.5万
  • 财政年份:
    2021
  • 负责人:
    Wang, Jiannan
  • 依托单位:
Crowdsourced Data Cleaning
  • 批准号:
    RGPIN-2016-05555
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.62万
  • 财政年份:
    2020
  • 负责人:
    Wang, Jiannan
  • 依托单位:
Entity augmentation and data cleaning for machine learning
  • 批准号:
    508081-2016
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $4.37万
  • 财政年份:
    2019
  • 负责人:
    Wang, Jiannan
  • 依托单位:
海外基金