Connecting Vision and Language in Instructional Videos through Human Gaze
Connecting Vision and Language in Instructional Videos through Human Gaze
批准号:
22K17905
负责人:
黄 逸飛
金额:
$2.91万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Early-Career Scientists
财政年份:
2022
资助国家:
日本
项目状态:
已结题
起止时间:
2022-04-01 至 2023-03-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The project unfortunately only last for half a year, however fruitful outcomes have still been made. In the research plan, we originally proposed to collect the world’s first instructional video dataset with synchronized language and gaze. Since the purchase of wearable equipment has not been ready due to the low productivity casued by Covid, in this fiscal year we collected a preliminary dataset using a similar setting, i.e., we recorded user's eye-tracking data on a fixed screen, containing a person’s gaze when he is watching and describing the video. In this way we could still analyze the dense connection between video, gaze and language. In total we have collected 2.5 hours of multi-modal data on both the video datasets. The audio data were collected in 15 different languages. We are still working on the publication of this dataset. While the total amount of collected data is limited, it is a very good start.Another part of research plan is to develop vision language algorithm aiming at publishing top-tier academic papers. Addressing the problem that we do not have enough data for model training at the moment, I developed few-shot based action recognition algorithm, which can significantly improve the model’s ability when given only less than 5 annotated samples. The paper has been accepted by top tier academic conference ECCV 2022.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.48550/arxiv.2207.05515
发表时间:
2022-07
期刊:
影响因子:
--
作者:
[Lijin Yang;Yifei Huang;Y. Sato]
通讯作者:
Lijin Yang;Yifei Huang;Y. Sato
Few-shot Action Recognition
少镜头动作识别
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
海外基金