课题基金 / 基金详情

Connecting Vision and Language in Instructional Videos through Human Gaze

Connecting Vision and Language in Instructional Videos through Human Gaze
通过人类凝视连接教学视频中的视觉和语言
批准号:
22K17905
负责人:
黄 逸飛
金额:
$2.91万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Early-Career Scientists
财政年份:
2022
资助国家:
日本
项目状态:
已结题
起止时间:
2022-04-01 至 2023-03-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
虽然项目只进行了半年,但还是取得了丰硕的成果。在研究计划中,我们最初提出要收集世界上第一个同步语言和凝视的教学视频数据集。由于Covid造成的低生产率,可穿戴设备的购买还没有准备好,在本财政年度,我们使用类似的设置收集了一个初步数据集,即我们在固定屏幕上记录用户的眼球追踪数据,其中包含一个人在观看和描述视频时的目光。通过这种方式,我们仍然可以分析视频、凝视和语言之间的紧密联系。总的来说,我们在两个视频数据集上收集了2.5小时的多模态数据。音频数据以15种不同的语言收集。我们仍在努力发布这个数据集。虽然收集的数据总量有限,但这是一个很好的开始。研究计划的另一部分是开发以发表顶级学术论文为目标的视觉语言算法。针对我们目前没有足够的数据进行模型训练的问题,我开发了基于few-shot的动作识别算法,该算法在给定少于5个带注释的样本时可以显著提高模型的能力。论文已被国际顶级学术会议ECCV 2022录用。
英文摘要
The project unfortunately only last for half a year, however fruitful outcomes have still been made. In the research plan, we originally proposed to collect the world’s first instructional video dataset with synchronized language and gaze. Since the purchase of wearable equipment has not been ready due to the low productivity casued by Covid, in this fiscal year we collected a preliminary dataset using a similar setting, i.e., we recorded user's eye-tracking data on a fixed screen, containing a person’s gaze when he is watching and describing the video. In this way we could still analyze the dense connection between video, gaze and language. In total we have collected 2.5 hours of multi-modal data on both the video datasets. The audio data were collected in 15 different languages. We are still working on the publication of this dataset. While the total amount of collected data is limited, it is a very good start.Another part of research plan is to develop vision language algorithm aiming at publishing top-tier academic papers. Addressing the problem that we do not have enough data for model training at the moment, I developed few-shot based action recognition algorithm, which can significantly improve the model’s ability when given only less than 5 annotated samples. The paper has been accepted by top tier academic conference ECCV 2022.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.48550/arxiv.2207.05515
发表时间: 2022-07
期刊:
影响因子: --
作者: [Lijin Yang;Yifei Huang;Y. Sato]
通讯作者: Lijin Yang;Yifei Huang;Y. Sato
Few-shot Action Recognition
少镜头动作识别
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
海外基金