课题基金 / 基金详情

Video Understanding and Retrieval with Complex Natural Language Queries

Video Understanding and Retrieval with Complex Natural Language Queries
通过复杂的自然语言查询进行视频理解和检索
批准号:
RGPIN-2015-04529
负责人:
Fidler, Sanja
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2016
资助国家:
加拿大
项目状态:
已结题
起止时间:
2016-01-01 至 2017-12-31

项目摘要

项目成果

Fidler, Sanja的其他基金

相似基金

相关文献

中文摘要
翻译
这项提议的目标是开发集成的语言和视觉解决方案,以实现日常交互中人类与自主系统之间的自然交流,以及人类与大规模视觉内容之间的自然交流,实现自然文本查询的智能浏览。语言是连接高层语义概念和更低层视觉感知的重要纽带。一个成功的机器人平台需要能够理解视觉世界和人类的(语言)指令,并以自然的方式将其理解传回给用户。我工作的主要兴趣领域是老年人和视障人士的语言辅助视觉感知解决方案。这是一个影响很大的研究领域,目前还没有好的解决方案。 该提案的重点将放在语言和视觉融合的三个不同方面,作为大型语言辅助视觉解决方案的一部分。第一个工作将是根据对象和活动对视觉场景进行有效的建模。我将继续我目前在对象检测方面的工作,我计划将多模式信息,如外观、语义分割、上下文信息、运动以及深度纳入深度学习框架。我希望这项工作将大大提高静态想象数据和视频数据的视觉识别性能。我的第二个努力将是将自然描述性文本的语义内容与视觉内容联系起来。我计划探索两个非常重要的设置,一个用于在线视觉检索,另一个用于离线大规模视觉内容浏览。在第一个设置中,人在直播视频流中向系统传达系统需要注意的描述性指令,并及时通知用户。这对于视障人士的解决方案尤其重要,其中自动系统需要充当用户的活动眼睛。在该提案将处理的相关设置中,用户给出复杂的描述性查询,并且系统需要从大量视频集合中检索最相关的视频剪辑。这对于分析大量记录的数据、了解用户特定的兴趣以及快速浏览Web规模的可视内容都很重要。这与现有搜索引擎只能处理一两个文本标签形成了鲜明对比。 这个方案的第三个目标是开发一个系统,能够通过语言描述他们对视觉场景的内部语义理解。虽然文本生成的解决方案已经存在,但没有一个解决方案试图与我们人类描述世界的结构和突出度相匹配。然而,如果一个自动系统要以一种自然的方式与人类交流,这是至关重要的。向用户描述对世界的视觉理解将关闭人与机器人交互的循环,并将使新一代机器人系统成为可能。
英文摘要
The goal of this proposal is to develop integrated language and vision solutions to enable natural communication between humans and autonomous systems in everyday interaction, as well as between humans and large-scale visual content, enabling intelligent browsing with natural textual queries. Language is an important link between high level semantic concepts and more low level visual perception. A successful robotic platform needs to be able to understand both the visual world and the human's (lingual) instructions and communicate its understanding back to the user in a natural way. The main domain of interest of my work are solutions for language assisted visual perception for the elderly and visually impaired. This is a high-impact area of research where no good solutions exist yet. The focus of the proposal will be on three different aspects of language and vision integration, as part of a large language assisted vision solution. The first effort will be on efficient modeling of visual scenes both in term of objects as well as activities. I will continue my current work on object detection, where I plan to incorporate multi-modal information such as appearance, semantic segmentation, contextual information, motion as well as depth into a deep learning framework. I expect this work to significantly push forward visual recognition performance in both static imaginary as well as video data. My second effort will be on linking the semantic content of natural descriptive text with visual content. I plan to explore two very important settings, one for online visual retrieval and one for offline large-scale visual content browsing. In the first setting, the human conveys to the system a descriptive instruction that the system needs to pay attention to in a live video stream, and inform the user in time. This is particularly important for a solution for the visually impaired, where the automatic system needs to serve as active eyes of the user. In a related setting the proposal will tackle, the user gives a complex descriptive query, and the system needs to retrieve the most relevant video clips from a large video collection. This is important to both analyze large amounts of recorded data, learn user-specific interests, as well as to quickly browse web-scale visual content. This is in contrast to existing search engines which can only handle one or two textual tags. The third goal in this proposal is to develop a system that can describe their internal semantic understanding of the visual scene via language. While solutions for text generation exist, none of them tries to match the structure and saliency with which us, humans, describe the world. This is, however, of critical importance, if an automatic system is to communicate with the human in a natural way. Describing the visual understanding of the world back to the user closes the loop of human-robot interaction and would enable a new generation of robotic systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Video Understanding and Retrieval with Complex Natural Language Queries
  • 批准号:
    RGPIN-2015-04529
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2021
  • 负责人:
    Fidler, Sanja
  • 依托单位:
Video Understanding and Retrieval with Complex Natural Language Queries
  • 批准号:
    RGPIN-2015-04529
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2020
  • 负责人:
    Fidler, Sanja
  • 依托单位:
Video Understanding and Retrieval with Complex Natural Language Queries
  • 批准号:
    RGPIN-2015-04529
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2019
  • 负责人:
    Fidler, Sanja
  • 依托单位:
Video Understanding and Retrieval with Complex Natural Language Queries
  • 批准号:
    RGPIN-2015-04529
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2018
  • 负责人:
    Fidler, Sanja
  • 依托单位:
国内基金
海外基金
Navigating Sustainability: Understanding Environm ent,Social and Governanc e Challenges and Solution s for Chinese Enterprises in Pakistan's CPEC Framew ork
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    Noshaba Aziz
  • 依托单位:
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
Understanding complicated gravitational physics by simple two-shell systems
  • 批准号:
    12005059
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    国分隆文
  • 依托单位: