课题基金 / 基金详情

Towards Open-world Semi-supervised learning

Towards Open-world Semi-supervised learning
走向开放世界的半监督学习
批准号:
2766068
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
大量完全注释的数据是当前深度学习模型成功的主要组成部分之一。然而,在许多应用程序场景中,拥有大量人工提供的注释的假设是不现实的,因为收集这些注释的成本很高,而且需要专业知识。此外,在现实世界中,随着时间的推移,新的概念和类别可能会出现。因此,为每个可能的概念收集人工注释通常是不可行的。这个项目将专注于如何建立自主代理的问题,这些代理可以在训练时不提供人类监督的情况下自动推理新类别。目前的半监督学习方法是有限的,因为它们要求所有类别至少有一个例子有人工注释。因此,在这个开放世界中,很难直接采用以前的半监督学习方法。最近出现了一个新的领域,称为新类别发现(NCD),它与这个项目密切相关,它关注的问题是如何利用现有标记数据集中的知识在大型未标记数据集中发现新类别。然而,NCD方法假设未标记集合中的所有类别都是新颖的,并且只关注这些新类别的性能。在现实世界中,能够识别以前看到的类别和发现新的类别同样重要。在这项工作中,我们的目标是开发开放世界半监督学习的算法和解决方案。开发的算法不仅能够使用标记和未标记的数据识别以前看到的类别,而且还能够从未标记的数据中发现新的类别。在成功开发这些新方法后,我们还旨在设计可解释的方法,以传达“为什么”模型认为一组示例是新颖的,以及新发现的新类别与以前见过的类别“如何”不同。这些建议的可解释方法将为大型未标记图像集提供新的见解,并将当前的深度学习方法推进到更现实的开放世界环境中。在这个项目中,我们将专注于以下目标:(1)分类器学习:学习使用标记和未标记的图像对已见和新类别进行分类。(2)类别数估计:类别数估计算法应能高效运行,并与分类器学习过程同步进行。(3)解释发现的类别:可解释的模型输出,以便用户知道为什么一些图像被预测形成一个新的类别。为了设计这种新的开放世界半监督学习环境的新方法,我们旨在从经典的无监督聚类方法和前沿的深度学习方法中汲取灵感:(1)层次聚类,这些方法具有自动合并数据点以自适应形成类别层次的优势,通过使用神经网络学习相似度量函数。我们可以利用它来执行分层聚类,并能够学习分类器,同时估计新类别的数量。(2)利用最先进的视觉语言模型,我们还旨在利用视觉语言学习的最新进展来设计自动生成文本解释的方法,以解释模型“为什么”认为一组示例是新颖的,以及新类别与人类标记数据集中的“看到”类别“如何”不同。我们将把这些方法应用到现实世界的任务中,以更好地理解目前的实际挑战,这将为开发更健壮的模型提供信息。
英文摘要
Large amounts of fully annotated data is one of the major components responsible for the success of current deep-learning models. However, in many application scenarios, this assumption of having an extensive collection of human provided annotations is not realistic, as gathering those annotations can be costly and require expert knowledge. Additionally in real-world settings, new concepts and categories may emerge over time. Thus it is often not feasible to gather human annotations for every possible concept. This project will focus on the question of how to build autonomous agents that can automatically reason about novel categories for which no human supervision is provided at training time.Current semi-supervised learning methods are limited in that they require all categories to have human annotations for at least one example. As a result it can be hard to directly adopt previous semi-supervised learning methods in this open-world regime. Recently a new area called Novel Category Discovery (NCD) has emerged, which is closely related to this project, which focuses on the problem of how to discover novel categories within a large unlabelled dataset by leveraging knowledge from an existing labelled dataset. However, NCD methods assume that all the categories in the unlabelled set are novel, and only focus on the performance of those novel categories.In a real-world scenario, it is of equal importance to be able to recognise both previously seen categories as well as discover novel ones. In this work, we aim to develop algorithms and solutions for open-world semi-supervised learning. The developed algorithms will not only be able to recognise previously seen categories using labelled and unlabelled data, but also be able to discover novel categories from unlabelled data. Upon successful development of these new methods, we also aim to design interpretable methods that can convey 'why' the model thinks a set of examples are novel and `how' a newly discovered novel category differs from previously seen ones. These proposed interpretable methods will give new insight into large unlabeled image collections and will advance current deep-learning approaches into a more realistic open-world setting. In this project, we will focus on the following goals: (1) Classifier learning: Learning to classify both seen and novel categories using labelled and unlabelled images.(2) Category number estimation: The category number estimation algorithm should be able to run efficiently and be performed simultaneously with the classifier learning process.(3) Interpreting discovered categories: Interpretable model outputs so that the user knows why some images are predicted to form a novel category.To design novel methods for this new open-world semi-supervised learning setting, we aim to draw inspiration from classic unsupervised clustering methods and cutting-edge deep learning methods:(1) Hierarchical clustering, these methods have the advantage of automatically merging data points to adaptively form a hierarchy of categories, by learning a similarity measuring function using neural networks. We can leverage it to perform hierarchical clustering and to be able to learn the classifier and estimate the number of novel categories at the same time.(2) Leveraging state-of-the-art visual-language models, we also aim to use recent advancements in visual-language learning to design methods for automatically generating a textual explanation for interpreting 'why' the model thinks a set of examples are novel, and 'how' the novel categories differ from the 'seen' categories from the human labelled dataset.We will apply these methods to real-world tasks to better understand the practical challenges present which will inform the development of more robust models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
精子发生中mRNA下游开放阅读框(downstream Open Reading Frame,dORF)的功能研究
  • 批准号:
    --
  • 项目类别:
    面上项目
  • 资助金额:
    54万元
  • 批准年份:
    2022
  • 负责人:
    刘明兮
  • 依托单位:
基于升阶谱方法和Open CASCADE的高阶网格自动生成技术研究
  • 批准号:
    11972004
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2019
  • 负责人:
    刘波
  • 依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
  • 批准号:
    61373035
  • 项目类别:
    面上项目
  • 资助金额:
    77.0万元
  • 批准年份:
    2013
  • 负责人:
    冯志勇
  • 依托单位:
变分与拓扑方法和Schrodinger方程中的Open 问题
  • 批准号:
    10871109
  • 项目类别:
    面上项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2008
  • 负责人:
    邹文明
  • 依托单位: