课题基金 / 基金详情

Utilizing unlabeled data for machine learning tasks - theoretical analysis

Utilizing unlabeled data for machine learning tasks - theoretical analysis
利用未标记数据进行机器学习任务 - 理论分析
批准号:
RGPIN-2015-04654
负责人:
BenDavid, Shai
金额:
$2.62万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

BenDavid, Shai的其他基金

相似基金

相关文献

中文摘要
翻译
主流机器学习工具依赖于人类标注的训练数据的可用性。如今,在所谓的“大数据”时代,机器学习的应用程序可以访问大量未注释(又称未标记)的数据。因此,人们对设计机器学习工具产生了巨大的兴趣,这些工具可以利用如此大的原始、未注释的数据池,以减少对学习过程中人为干预的需求。最近出现的各种机器学习范例都解决了这个问题。这些方法包括聚类、主动学习、半监督学习、领域适应和迁移学习,以及向“弱教师”学习(比如通过众包获得的监督)。为了应对这种情况,学习实践者已经开发了启发式,虽然在实践中显然相当有效,但现有的数学分析并不支持。在过去的几十年里,机器学习为理论分析对实际应用发展的影响提供了一个响亮的证明。支持向量机(Support-Vector-Machines)、决策树(Decision-Trees)和助推(Boosting)等算法范例从理论模型发展成为流行且广泛适用的软件包。机器学习理论分析的成功是否可以扩展到现代大数据最小人为干预场景?提出的研究旨在通过为那些新兴的机器学习和数据挖掘范式建立数学支持,为这些发展提供基础。***我的团队最近在这个方向上采取的一些举措包括一个程序,旨在提供工具,指导希望聚类大数据集的用户如何选择适当的聚类算法和参数设置。这些选择对于集群应用程序的成功至关重要,但到目前为止都是以一种临时的方式完成的。我的学生Margareta Ackerman的博士论文在这个方向上迈出了第一步,但仍然有很大的挑战需要克服,无论是在开发这样的工具方面,还是在提高数据挖掘社区对聚类任务和用于解决这些任务的算法之间匹配的重要性的认识方面。另一个项目解决了利用“弱监管”的任务。由于使用众包的日益普及,新手主管使用注释来帮助收集分类预测任务的训练数据最近引起了研究的关注。与目前在这个方向上的许多研究相反,我们考虑了弱监督与人类监督合作的场景,旨在减少(而不是消除)对专家的呼叫。我们正在开发新的数学模型,以放大那些由新手生成的标签不可信、需要由人类专家仔细审查的实例
英文摘要
Mainstream machine learning tools depend on the availability of human annotated training data. Nowaday, in what is called the ``big data" era, applications of machine learning have access to very large amounts of un-annotated (a.k.a. unlabeled) data. Consequently, there is vast interest in designing machine learning tools that can utilize such big pools of raw, unannotated data to reduce the need for human intervention in the learning process. Various recently arising machine learning paradigms address this issue. These include Clustering, Active Learning, Semi-Supervised Learning, Domain Adaptation and Transfer Learning, as well as Learning from ``Weak Teachers" (like supervision obtained via crowdsourcing). To cope with such scenarios, learning practitioners have developed heuristics that, while apparently working reasonably well in practice, are not supported by existing mathematical analyses. ***In the past couple of decades, machine learning provided a resounding demonstration of the impact of theoretical analysis on the development of practical applications. Algorithmic paradigms like Support-Vector-Machines, Decision-Trees and Boosting grew from theoretical models into popular and vastly applicable software packages. Can the success of theoretical analysis of machine learning be extended to the modern big-data-minimal-human-intervention scenarios? The proposed research aims to provide basis for such developments by building mathematical support for those emerging machine learning and data mining paradigms. *** Some examples of recent initiatives taken by my team in that direction include a program that aims to provide tools for guiding users that wish to cluster big data sets on how to choose appropriate clustering algorithms and parameter settings. Such choices are critical to the success of clustering applications, and yet have so far been done in an ad hoc fashion. The PhD thesis of my student Margareta Ackerman took first steps in this direction but there are still big challenges to overcome, both in terms of developing such tools, and in terms of raising the awareness of the data mining community to the significance of the matching between clustering tasks and the algorithms employed to address them. Another project addresses the task of utilizing ``weak supervision".  The use of annotation by novice supervisors to help collect training data for classification prediction tasks has been drawing research attention recently due to the growing popularity of using crowdsourcing. In contrast with much of the current research in this direction, we consider the scenario in which the weak supervision is used in cooperation with human supervision, aiming to reduce  (rather than eliminate) calls to the expert. We are developing novel mathematical models to zoom in on the instances for which novice-generated labels cannot be trusted and need to be scrutinized by a human expert.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Machine Learning Beyond Prediction - Extracting Insights and Guiding Actions
  • 批准号:
    RGPIN-2020-04333
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2022
  • 负责人:
    BenDavid, Shai
  • 依托单位:
Machine Learning Beyond Prediction - Extracting Insights and Guiding Actions
  • 批准号:
    RGPIN-2020-04333
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2021
  • 负责人:
    BenDavid, Shai
  • 依托单位:
Machine Learning Beyond Prediction - Extracting Insights and Guiding Actions
  • 批准号:
    RGPIN-2020-04333
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2020
  • 负责人:
    BenDavid, Shai
  • 依托单位:
Utilizing unlabeled data for machine learning tasks - theoretical analysis
  • 批准号:
    RGPIN-2015-04654
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.62万
  • 财政年份:
    2019
  • 负责人:
    BenDavid, Shai
  • 依托单位:
海外基金