Vision and language cross-modal for training conditional GANs with long-tail data.
Vision and language cross-modal for training conditional GANs with long-tail data.
批准号:
22K17947
负责人:
ヴォ ミンデュク
金额:
$1.66万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Early-Career Scientists
财政年份:
2022
资助国家:
日本
项目状态:
已结题
起止时间:
2022-04-01 至 2024-03-31
中文摘要
我们学习视觉和语言空间之间的交叉模态。本文的主要研究成果有三:1.从开放词典Wiktionary中收集了一组对象的名称和定义,并使用预先训练好的BERT模型嵌入定义。我们将这种外部知识的图像字幕模型,优于其他方法在新的对象字幕任务。它发表在CVPR 2022.2上。我们通过使用翻转和非翻转非饱和损耗,提出了一种新的GAN训练方案。3.我们创建了一个新的故事评估数据集,包括通过Reddit网站和众包注释过程收集的10万个故事排名数据和46k方面评级和推理。它发表在EMNLP 2022上。
英文摘要
We learn the cross-modality between vision and language spaces. We obtained three achievements:1.We collected a set of objects' names and definitions from the open dictionary Wiktionary and used the pre-trained BERT model to embed the definitions. We incorporated this external knowledge into an image captioning model, outperforming other methods in novel object captioning task. It was published at CVPR 2022.2.We proposed a new training scheme for GANs by using flipped and non-flipped non-saturating losses. It was published in the IEEE Access journal (IF 3.476).3.We created a new dataset for story evaluation, consisting of 100k story ranking data and 46k aspect rating and reasoning collected through the Reddit website and crowd-sourcing annotation process. It was published at EMNLP 2022.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1109/access.2022.3210130
发表时间:
2022
期刊:
IEEE Access
影响因子:
3.9
作者:
[Rui Yang;Duc Minh Vo;Hideki Nakayama]
通讯作者:
Rui Yang;Duc Minh Vo;Hideki Nakayama
NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External Knowledge
NOC-REK:从外部知识检索词汇的新颖对象描述
DOI:
10.1109/cvpr52688.2022.01747
发表时间:
2022
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
作者:
[Duc Minh Vo, Hong Chen, Akihiro Sugimoto, Hideki Nakayama]
通讯作者:
Hideki Nakayama
DOI:
10.48550/arxiv.2210.08459
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
作者:
[Hong Chen;Duc Minh Vo;Hiroya Takamura;Yusuke Miyao;Hideki Nakayama]
通讯作者:
Hong Chen;Duc Minh Vo;Hiroya Takamura;Yusuke Miyao;Hideki Nakayama
Unifying Object Detection and Image Captioning using Vision-Language Knowledge Base for Open-World Comprehension
-
批准号:24K20830
-
项目类别:Grant-in-Aid for Early-Career Scientists
-
资助金额:$3.0万
-
财政年份:2024
-
负责人:ヴォ ミンデュク
-
依托单位: