课题基金 / 基金详情

Video Understanding and Retrieval with Complex Natural Language Queries

Video Understanding and Retrieval with Complex Natural Language Queries
通过复杂的自然语言查询进行视频理解和检索
批准号:
RGPIN-2015-04529
负责人:
Fidler, Sanja
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Fidler, Sanja的其他基金

相似基金

相关文献

中文摘要
翻译
该提案的目标是开发集成的语言和视觉解决方案,以实现人类与自治系统之间以及人类与大规模视觉内容之间的日常交互中的自然通信,从而实现具有自然文本查询的智能浏览。语言是高级语义概念与低级视觉感知之间的重要纽带。一个成功的机器人平台需要能够理解视觉世界和人类(语言)指令,并以自然的方式将其理解传达给用户。我的主要工作兴趣领域是为老年人和视障人士提供语言辅助视觉感知的解决方案。这是一个影响很大的研究领域,目前还没有好的解决方案。***该提案的重点将放在语言和视觉整合的三个不同方面,作为一个大型语言辅助视觉解决方案的一部分。第一个努力将是在对象和活动方面对视觉场景的有效建模。我将继续我目前在物体检测方面的工作,在那里我计划将多模态信息,如外观,语义分割,上下文信息,运动以及深度纳入深度学习框架。我希望这项工作能够显著地推动静态想象和视频数据中的视觉识别性能。我的第二个努力将是连接自然描述性文本的语义内容与视觉内容。我计划探索两个非常重要的设置,一个用于在线视觉检索,一个用于离线大规模视觉内容浏览。在第一种设置中,人在实时视频流中向系统传达系统需要注意的描述性指令,并及时通知用户。这对于视障人士的解决方案尤其重要,因为自动系统需要充当用户的主动眼睛。在一个相关的设置中,建议将处理,用户给出一个复杂的描述性查询,系统需要从一个大的视频集合中检索最相关的视频剪辑。这对于分析大量记录数据、了解用户特定兴趣以及快速浏览网络规模的视觉内容都很重要。这与现有的只能处理一个或两个文本标签的搜索引擎形成了对比。***本提案的第三个目标是开发一个系统,可以通过语言描述他们对视觉场景的内部语义理解。虽然存在文本生成的解决方案,但它们都没有试图与我们人类描述世界的结构和显著性相匹配。然而,如果一个自动系统要以一种自然的方式与人类交流,这是至关重要的。向用户描述对世界的视觉理解闭合了人机交互的循环,并将使新一代机器人系统成为可能
英文摘要
The goal of this proposal is to develop integrated language and vision solutions to enable natural communication between humans and autonomous systems in everyday interaction, as well as between humans and large-scale visual content, enabling intelligent browsing with natural textual queries. Language is an important link between high level semantic concepts and more low level visual perception. A successful robotic platform needs to be able to understand both the visual world and the human's (lingual) instructions and communicate its understanding back to the user in a natural way. The main domain of interest of my work are solutions for language assisted visual perception for the elderly and visually impaired. This is a high-impact area of research where no good solutions exist yet.***The focus of the proposal will be on three different aspects of language and vision integration, as part of a large language assisted vision solution. The first effort will be on efficient modeling of visual scenes both in term of objects as well as activities. I will continue my current work on object detection, where I plan to incorporate multi-modal information such as appearance, semantic segmentation, contextual information, motion as well as depth into a deep learning framework. I expect this work to significantly push forward visual recognition performance in both static imaginary as well as video data. My second effort will be on linking the semantic content of natural descriptive text with visual content. I plan to explore two very important settings, one for online visual retrieval and one for offline large-scale visual content browsing. In the first setting, the human conveys to the system a descriptive instruction that the system needs to pay attention to in a live video stream, and inform the user in time. This is particularly important for a solution for the visually impaired, where the automatic system needs to serve as active eyes of the user. In a related setting the proposal will tackle, the user gives a complex descriptive query, and the system needs to retrieve the most relevant video clips from a large video collection. This is important to both analyze large amounts of recorded data, learn user-specific interests, as well as to quickly browse web-scale visual content. This is in contrast to existing search engines which can only handle one or two textual tags.***The third goal in this proposal is to develop a system that can describe their internal semantic understanding of the visual scene via language. While solutions for text generation exist, none of them tries to match the structure and saliency with which us, humans, describe the world. This is, however, of critical importance, if an automatic system is to communicate with the human in a natural way. Describing the visual understanding of the world back to the user closes the loop of human-robot interaction and would enable a new generation of robotic systems.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Video Understanding and Retrieval with Complex Natural Language Queries
  • 批准号:
    RGPIN-2015-04529
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2021
  • 负责人:
    Fidler, Sanja
  • 依托单位:
Video Understanding and Retrieval with Complex Natural Language Queries
  • 批准号:
    RGPIN-2015-04529
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2020
  • 负责人:
    Fidler, Sanja
  • 依托单位:
Video Understanding and Retrieval with Complex Natural Language Queries
  • 批准号:
    RGPIN-2015-04529
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2019
  • 负责人:
    Fidler, Sanja
  • 依托单位:
Video Understanding and Retrieval with Complex Natural Language Queries
  • 批准号:
    RGPIN-2015-04529
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2017
  • 负责人:
    Fidler, Sanja
  • 依托单位:
国内基金
海外基金
Navigating Sustainability: Understanding Environm ent,Social and Governanc e Challenges and Solution s for Chinese Enterprises in Pakistan's CPEC Framew ork
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    Noshaba Aziz
  • 依托单位:
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
Understanding complicated gravitational physics by simple two-shell systems
  • 批准号:
    12005059
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    国分隆文
  • 依托单位: