课题基金 / 基金详情

RI: Small: Learning Fine-Grained Instructions from Uncurated Complex Activity Videos

RI: Small: Learning Fine-Grained Instructions from Uncurated Complex Activity Videos
RI:小型:从未经策划的复杂活动视频中学习细粒度的指令
批准号:
2115110
负责人:
Ehsan Elhamifar
金额:
$49.77万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-10-01 至 2024-09-30

项目摘要

项目成果

Ehsan Elhamifar的其他基金

相似基金

相关文献

中文摘要
翻译
人类有一种非凡的能力,可以通过观察别人的行为并听从他们的指示来学习完成复杂的任务。将这种能力引入机器对人工智能的发展具有深远的影响,例如设计智能助手和机器人,可以通过挖掘教学和日常活动视频来学习执行或指导人类完成任务。尽管最近取得了一些进展,但视频和活动理解方法仍然面临着重大挑战,即如何将未经修剪的复杂活动的原始长视频转换为详细而准确的说明。这些问题包括视频中指令的巨大外观和运动变化,从长视频中收集密集的时间视频注释的高成本,缺乏一种系统的方法来整合不同类型的可用噪声但廉价的标签以进行有效的学习,以及难以生成长期的未来指令。该项目研究了一个全面的数学框架,用于从未经修剪的长而复杂的活动视频中学习详细而准确的指令,克服了上述挑战。该研究项目伴随着一项综合教育和推广计划,该计划包括通过东北大学的青年学者计划指导高中生和本科生,并将项目成果整合到本科生和研究生课程中。该项目将公开发布实现开发算法的开源软件。该项目开发了新的无监督和自监督任务分割和子任务(指令步骤)定位方法,通过研究任务的多流形模型,同时学习和发现视频流形之间的关联,同时结合任务约束和先验。开发的框架允许处理跨视频子任务的大型外观和动作变化,并允许利用其他模式,如视频叙述和音频。研究小组将开发一个统一的弱监督视觉接地框架,该框架基于深度神经网络,从不同类型的可用廉价噪声弱标签中学习,处理分布尾部的子任务,并从当前观察中生成未来的指令。此外,该团队将研究一种新的概率深度学习框架,该框架具有与子任务、语法和任务预测相对应的分层连接模块,允许集成所有类型的弱标签并生成可信的未来子任务序列。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Humans have the remarkable ability of learning to perform complex tasks by watching others performing them and following their instructions. Bringing this capability to machines has far reaching impact on the advancement of the artificial intelligence with important applications, such as designing intelligent assistants and robots that can learn to perform or guide humans through tasks by mining instructional and everyday activity videos. Despite recent advances, there are major challenges facing video and activity understanding methods to convert raw untrimmed long videos of complex activities into detailed and accurate instructions. These include large appearance and motion variations of instructions across videos, high cost of gathering dense temporal video annotations from long videos, lack of a systematic way of integrating different types of available noisy yet inexpensive labels for effective learning and difficulty of generating long-range future instructions. This project investigates a comprehensive mathematical framework for learning detailed and accurate instructions from untrimmed long complex activity videos, overcoming the aforementioned challenges. The research project is accompanied with an integrated education and outreach plan, which involves mentoring high school and undergraduate students through the Northeastern's Young Scholar Program and integrating the results of the project into the undergraduate and graduate classes. The project will publicly release an open-source software implementing the developed algorithms.This project develops new unsupervised and self-supervised task segmentation and subtask (instruction step) localization methods, by investigating a multi-manifold model for tasks and simultaneously learning and finding associations between manifolds across videos while incorporating task constraints and priors. The developed framework allows for handling large appearance and motion variations of subtasks across videos and allows for leveraging other modalities, such as video narrations and audio. The research team will develop a unified weakly-supervised visual grounding framework based on deep neural networks that learns from different types of available inexpensive noisy weak labels, handles subtasks at the distribution tail and generates future instructions from current observations. Furthermore, the team will investigate a new probabilistic deep learning framework with hierarchically connected modules corresponding to subtask, grammar and task prediction, allowing to integrate all types of weak labels and to generate plausible future subtask sequences.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cvpr52688.2022.00334
发表时间: 2022-06
期刊: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者: [Yuhan Shen;Ehsan Elhamifar]
通讯作者: Yuhan Shen;Ehsan Elhamifar
DOI: 10.1109/cvpr52688.2022.01928
发表时间: 2022-06
期刊: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者: [Zijia Lu;Ehsan Elhamifar]
通讯作者: Zijia Lu;Ehsan Elhamifar
Learning to Segment Actions from Visual and Language Instructions via Differentiable Weak Sequence Alignment
学习通过可微弱序列对齐从视觉和语言指令中分割动作
DOI: 10.1109/cvpr46437.2021.01002
发表时间: 2022
期刊: IEEE Conference on Computer Vision and Pattern Recognition
影响因子: --
作者: [Shen, Y., Wang, L., Elhamifar, E.]
通讯作者: Elhamifar, E.
CRII: RI: Towards a Comprehensive Dynamic Subset Selection Framework
  • 批准号:
    1657197
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.49万
  • 财政年份:
    2017
  • 负责人:
    Ehsan Elhamifar
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: