TVLT: Textless Vision-Language Transformer

TVLT: Textless Vision-Language Transformer
复制标题

DOI:
10.48550/arxiv.2209.14156
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Zineng Tang;Jaemin Cho;Yixin Nie;Mohit Bansal
Zineng Tang;Jaemin Cho;Yixin Nie;Mohit Bansal
中科院分区:
其他
文献类型:
--
作者:
Zineng Tang;Jaemin Cho;Yixin Nie;Mohit Bansal

文献摘要

被引文献

相似文献

在这项工作中,我们提出了无文本的视觉语言Transformer(TVLT),其中同质的Transformer块采取原始的视觉和音频输入的视觉和语言表示学习与最小的模态特定的设计,不使用文本特定的模块,如标记或自动语音识别(ASR)。TVLT通过重建连续视频帧和音频频谱图的掩蔽补丁(掩蔽自动编码)和对比建模来对齐视频和音频来训练。TVLT在各种多模态任务(如视觉问答、图像检索、视频检索和多模态情感分析)上的性能可与基于文本的同类产品相媲美,推理速度快28倍,参数仅为1/3。我们的研究结果表明,学习紧凑和有效的视觉语言表示从低级别的视觉和音频信号,而不假设文本的存在。我们的代码和检查点可在https://github.com/zinengtang/TVLT上找到
In this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representation learning with minimal modality-specific design, and do not use text-specific modules such as tokenization or automatic speech recognition (ASR). TVLT is trained by reconstructing masked patches of continuous video frames and audio spectrograms (masked autoencoding) and contrastive modeling to align video and audio. TVLT attains performance comparable to its text-based counterpart on various multimodal tasks, such as visual question answering, image retrieval, video retrieval, and multimodal sentiment analysis, with 28x faster inference speed and only 1/3 of the parameters. Our findings suggest the possibility of learning compact and efficient visual-linguistic representations from low-level visual and audio signals without assuming the prior existence of text. Our code and checkpoints are available at: https://github.com/zinengtang/TVLT