Web-Scale Semantic Image and Video Understanding
Web-Scale Semantic Image and Video Understanding
批准号:
RGPIN-2018-04657
负责人:
Sigal, Leonid
金额:
$2.99万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31
中文摘要
视觉识别是计算机视觉的一个子领域,其核心是建立能够自动和智能地识别、分类和理解图像/视频内容的算法。近年来,在强大的机器学习算法(例如深度学习)、大型标签数据集和新颖的问题定义的推动下,识别问题在准确性和范围方面取得了重大进展。然而,目前的方法缺乏允许准确、细粒度的详细场景理解的能力。大多数算法仅限于粗略的图像或视频级别的解释(例如,在1000个名词类别、200个动作类或视觉问题的18个候选答案中进行选择)。这样的理解是有用的,但同时在实现从媒体管理到自主导航的潜在应用的广度方面仍然非常有限。*很难量化视觉识别取得广泛成功所需的性能或规模水平。我相信下一个变革性的里程碑将是--在当前粗略模型的精确度下对场景的详细理解。这需要开发识别细粒度对象类别的能力,编号为100,000,以及对象和场景元素之间的时空(谓词)关系,计数为1,000。前者将赋予识别几乎每一个对象/名词的能力;后者将使情境语境推理对人工智能至关重要。我的长期研究目标是为大规模的视觉理解开发这样准确详细的模型;这些模型可以描述和定位对象和人,推理他们的空间和功能关系,他们的行为和相互作用。*这项建议在相应的研究思路中解决了实现这一目标的三个基本子挑战:*1.不断增长的细粒度类集合需要开发新的数据高效学习算法。随着要识别的类别变得更加具体,每个类别的数据量会减少(例如,有数百万辆汽车图像,但1957年的捷豹XKSS图像很少)。我们将在我们最近工作的基础上,开发出迄今为止唯一能够识别多达310,000个类别的方法。*2.超越了对孤立物体的识别,需要对空间和时间上将物体、人和场景元素联系起来的结构进行推理。将开发丰富灵活的结构化模型来实现这种推理。*3.为了缓解现有架构的黑箱性质,不适合决策关键任务,我们将开发支持可解释性和更像人类的内省推理的算法。*重要的是,该计划还将专注于将开发的算法应用于与媒体搜索/检索、增强现实和医学成像相关的特定识别问题。
英文摘要
Visual recognition is a sub-field of computer vision which centers on building algorithms that can automatically and intelligently recognize, catalog and understand image/video content. Significant progress in accuracy and scope has been made on recognition problems in recent years, driven by powerful machine learning algorithms (e.g., deep learning), large labeled datasets and novel problem definitions. However, current approaches lack capabilities that permit accurate fine-grained detailed scene understanding. Most algorithms are limited to coarse image- or video-level interpretations (e.g., choosing among 1,000 noun categories, 200 action classes, or 18 candidate answers for a visual question). Such understanding is useful, but at the same time is still very limiting in enabling the breadth of potential applications ranging from media curation to autonomous navigation. ******It is difficult to quantify what level of performance or scale is necessary for visual recognition to be broadly successful. I believe the next transformative milestone to be - detailed scene understanding at the accuracy of the current coarse models. This requires developing capabilities of recognizing fine-grained object categories, numbering in 100,000, and spatio-temporal (predicate) relations among the objects and elements of the scene, counting in 1,000. The former would give ability to recognize nearly every object/noun; the latter would enable situated contextual reasoning critically important for AI. My long term research objective is to develop such accurate detailed models for visual understanding at scale; models that can describe and localize objects and people, reason about their spatial and functional relationships, their actions and interactions. ******This proposal tackles three fundamental sub-challenges to achieving this objective in the corresponding research threads: ******1. The ever growing fine-grained set of classes requires development of novel data efficient learning algorithms. As the categories to recognize become more specific, the amount of data per category decreases (e.g., there are millions of car images, but few of 1957 Jaguar XKSS). We will build on our recent work where we developed the only method to date capable of recognizing up to 310,000 categories. ******2. Moving beyond recognition of isolated objects, requires reasoning about structures relating objects, people, and scene elements in space and time. Rich flexible structured models will be developed to enable such reasoning. ******3. To alleviate the black-box nature of existing architectures, not suitable for decision-critical tasks, we will develop algorithms that enable interpretability and more human-like introspective reasoning. ******Importantly, the program will also focus on applying the developed algorithms to specific recognition problems relevant for media search/retrieval, augmented reality and medical imaging.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Web-Scale Semantic Image and Video Understanding
-
批准号:RGPIN-2018-04657
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$5.97万
-
财政年份:2022
-
负责人:Sigal, Leonid
-
依托单位:
Canada Research Chair in Computer Vision and Machine Learning
-
批准号:CRC-2017-00182
-
项目类别:Canada Research Chairs
-
资助金额:$8.74万
-
财政年份:2022
-
负责人:Sigal, Leonid
-
依托单位:
Canada Research Chair In Computer Vision And Machine Learning
-
批准号:CRC-2017-00182
-
项目类别:Canada Research Chairs
-
资助金额:$8.74万
-
财政年份:2021
-
负责人:Sigal, Leonid
-
依托单位:
Web-Scale Semantic Image and Video Understanding
-
批准号:RGPIN-2018-04657
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.99万
-
财政年份:2021
-
负责人:Sigal, Leonid
-
依托单位:
Canada Research Chair in Computer Vision and Machine Learning
-
批准号:CRC-2017-00182
-
项目类别:Canada Research Chairs
-
资助金额:$8.74万
-
财政年份:2020
-
负责人:Sigal, Leonid
-
依托单位:
Web-Scale Semantic Image and Video Understanding
-
批准号:RGPIN-2018-04657
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.99万
-
财政年份:2020
-
负责人:Sigal, Leonid
-
依托单位:
Web-Scale Semantic Image and Video Understanding
-
批准号:RGPIN-2018-04657
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.99万
-
财政年份:2019
-
负责人:Sigal, Leonid
-
依托单位:
Web-Scale Semantic Image and Video Understanding
-
批准号:522579-2018
-
项目类别:Discovery Grants Program - Accelerator Supplements
-
资助金额:$5.83万
-
财政年份:2019
-
负责人:Sigal, Leonid
-
依托单位:
Canada Research Chair in Computer Vision and Machine Learning
-
批准号:CRC-2017-00182
-
项目类别:Canada Research Chairs
-
资助金额:$8.74万
-
财政年份:2019
-
负责人:Sigal, Leonid
-
依托单位:
Web-Scale Semantic Image and Video Understanding
-
批准号:522579-2018
-
项目类别:Discovery Grants Program - Accelerator Supplements
-
资助金额:$2.91万
-
财政年份:2018
-
负责人:Sigal, Leonid
-
依托单位:
Canada Research Chair in Computer Vision and Machine Learning
-
批准号:CRC-2017-00182
-
项目类别:Canada Research Chairs
-
资助金额:$8.74万
-
财政年份:2018
-
负责人:Sigal, Leonid
-
依托单位:
国内基金
海外基金
基于热量传递的传统固态发酵过程缩小(Scale-down)机理及调控
-
批准号:22108101
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:靳光远
-
依托单位:
基于Multi-Scale模型的轴流血泵瞬变流及空化机理研究
-
批准号:31600794
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2016
-
负责人:荆腾
-
依托单位:
针对Scale-Free网络的紧凑路由研究
-
批准号:60673168
-
项目类别:面上项目
-
资助金额:25.0万元
-
批准年份:2006
-
负责人:张国清
-
依托单位: