课题基金 / 基金详情

Foundations of Unsupervised and Weakly Supervised Learning

Foundations of Unsupervised and Weakly Supervised Learning
无监督和弱监督学习的基础
批准号:
RGPIN-2019-06018
负责人:
Ashtiani, Hassan
金额:
$1.68万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Ashtiani, Hassan的其他基金

相似基金

相关文献

中文摘要
翻译
能够从不断增长的数据来源中提取有用结构的计算工具的发展,正在改变分析或设计复杂系统的方式。随着标准机器学习方法工作的基本假设不再有效,在新的应用中使用机器学习的不断推动提出了新的理论和实践挑战。虽然实践者在这些情况下经常求助于特定于案例的解决方案作为第一道防线,但建立更通用和可靠的设计原则是至关重要的。我们的计划是解决一些相关的无监督(和弱监督)学习问题:我们的目标是从数学上表征这些环境下学习方法的成功。 无监督学习的重要性源于这样一个事实,即这些方法可以利用未注释的数据(这些数据在许多领域都很丰富)。非监督学习问题经常出现在科学和工程中(例如,在探索性数据分析、推荐系统、语音/文本/图像生成和医学成像中)。然而,从理论上看,与有监督学习相比,无监督学习的许多领域仍然不发达,启发式方法通常被采用,而没有提供有意义的用户级保证。我们旨在通过制定和分析一组无监督学习范例来解决这一缺陷,并提供可证明有效的方法(在计算和/或统计复杂性方面)来解决这些问题。 我们将朝着以下方向努力: 用于集群的系统化模型选择方案。集群方法在实践中被广泛使用,但对于“哪种集群方法适合我的用例?”这一基本问题。除了试错之外,很难找到既定的解决方案。为了解决这个问题,我们寻求以有原则的方式利用领域知识(例如,交互聚类或弱监管的聚类)。 分配的有效学习和测试。学习和测试未知分布(给定由其生成的样本)是统计学中的经典问题。然而,处理高维数据的需求带来了新的挑战。我们寻求开发不仅在统计上有效而且在计算上容易处理的方法。此外,在自然语言/语音生成等应用程序的激励下,我们研究了关于根本不同的距离度量(例如,对抗性距离)的分布学习/测试。 具有稀缺训练数据的有监督学习。在许多应用(例如,医疗诊断)中,标注训练数据是昂贵的,甚至是不可能的。因此,我们寻求开发在注释训练数据方面要求较低的解决方案(例如,使用无监督和弱监督学习)。我们还研究了学习的“有效样本复杂性”,特别是对于深度神经网络。
英文摘要
The development of computational tools that can extract useful structures from the ever-growing sources of data is transforming the way complex systems are analyzed or engineered. The constant push to use machine learning in new applications poses new theoretical and practical challenges, as the fundamental assumptions under which the standard machine learning methods work are no longer valid. While practitioners often resort to case-specific solutions in these situations as the first line of defense, it is essential to establish more generic and reliable design principles. Our plan is to address this for a number of related unsupervised (and weakly supervised) learning problems: we aim at mathematically characterizing the success of learning methods in those settings. The significance of unsupervised learning stems from the fact that these methods can utilize unannotated data (which is abundant in many domains). Unsupervised learning problems arise frequently in science and engineering (e.g., in exploratory data analysis, recommender systems, speech/text/image generation, and medical imaging). From the theoretical point of view, however, many areas of unsupervised learning are still under-developed compared to supervised learning, and heuristic methods are routinely adopted without offering meaningful user-level guarantees. We aim at addressing this shortcoming by formulating and analyzing a set of unsupervised learning paradigms, and providing provably efficient methods (in terms of computational and/or statistical complexities) for solving them. We will work on these directions: Systematic Model Selection Schemes for Clustering. Clustering methods are widely used in practice, but for the basic question of “which clustering method is good for my use case?” it is hard to find an established solution beyond trial and error. To address this, we seek to exploit domain knowledge in a principled way (e.g., interactive clustering or clustering with weak supervision). Efficient Learning and Testing of Distributions. Learning and testing an unknown distribution (given a sample generated from it) are classic problems in statistics. Fresh challenges, however, are arising from the need for handling high-dimensional data. We seek to develop methods that are not only statistically efficient but are also computationally tractable. Also, motivated by applications such as natural-language/speech generation, we study distribution learning/testing with respect to fundamentally different distance measures (e.g., adversarial distances). Supervised Learning with Scarce Training Data. In many applications (e.g., medical diagnosis) it is costly or even impossible to annotate the training data. Therefore, we seek to develop solutions that are less demanding in terms of annotated training data (using, e.g., unsupervised and weakly supervised learning). We also study the "effective sample complexity" of learning particularly for deep neural networks.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Foundations of Unsupervised and Weakly Supervised Learning
  • 批准号:
    RGPIN-2019-06018
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2022
  • 负责人:
    Ashtiani, Hassan
  • 依托单位:
Foundations of Unsupervised and Weakly Supervised Learning
  • 批准号:
    RGPIN-2019-06018
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2021
  • 负责人:
    Ashtiani, Hassan
  • 依托单位:
Foundations of Unsupervised and Weakly Supervised Learning
  • 批准号:
    RGPIN-2019-06018
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2019
  • 负责人:
    Ashtiani, Hassan
  • 依托单位:
Foundations of Unsupervised and Weakly Supervised Learning
  • 批准号:
    DGECR-2019-00016
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2019
  • 负责人:
    Ashtiani, Hassan
  • 依托单位:
海外基金