TAP: Text-Aware Pre-training for Text-VQA and Text-Caption

TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
复制标题

DOI:
10.1109/cvpr46437.2021.00864
复制
发表时间:
2020-12
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Zhengyuan Yang;Yijuan Lu;Jianfeng Wang;Xi Yin;D. Florêncio;Lijuan Wang;Cha Zhang;Lei Zhang
Zhengyuan Yang;Yijuan Lu;Jianfeng Wang;Xi Yin;D. Florêncio;Lijuan Wang;Cha Zhang;Lei Zhang
中科院分区:
其他
文献类型:
--
作者:
Zhengyuan Yang;Yijuan Lu;Jianfeng Wang;Xi Yin;D. Florêncio;Lijuan Wang;Cha Zhang;Lei Zhang

文献摘要

被引文献

相似文献

在本文中,我们提出了文本感知预训练(TAP)的文本VQA和文本标题任务。这两个任务的目标分别是阅读和理解图像中的场景文本,用于问答和图像字幕生成。传统的视觉语言预训练无法捕捉场景文本及其与视觉和文本模态的关系,而TAP在预训练期间明确地包含场景文本(从OCR引擎生成)。通过三个预训练任务,包括掩码语言建模(MLM),图像-文本(对比)匹配(ITM)和相对(空间)位置预测(RPP),使用场景文本进行预训练可以有效地帮助模型在三种模态(文本词,视觉对象和场景文本)之间学习更好的对齐表示。由于这种对齐的表示学习,即使在相同的下游任务数据集上进行了预训练,与非TAP基线相比,TAP已经将TextVQA数据集的绝对准确率提高了+5:4%。为了进一步提高性能,我们建立了一个大规模的场景文本相关的imagetext数据集的基础上的概念字幕数据集,命名为OCR-CC,其中包含1:4万的图像与场景文本。在这个OCR-CC数据集上进行了预训练,我们的方法在多个任务上的表现远远优于现有技术,即,TextVQA的准确率为+8:3%,ST-VQA的准确率为+8:6%,TextCaps的CIDEr评分为+10:2。
In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answering and image caption generation, respectively. In contrast to conventional vision-language pretraining that fails to capture scene text and its relationship with the visual and text modalities, TAP explicitly incorporates scene text (generated from OCR engines) during pretraining. With three pre-training tasks, including masked language modeling (MLM), image-text (contrastive) matching (ITM), and relative (spatial) position prediction (RPP), pre-training with scene text effectively helps the model learn a better aligned representation among the three modalities: text word, visual object, and scene text. Due to this aligned representation learning, even pre-trained on the same downstream task dataset, TAP already boosts the absolute accuracy on the TextVQA dataset by +5:4%, compared with a non-TAP baseline. To further improve the performance, we build a large-scale scene text-related imagetext dataset based on the Conceptual Caption dataset, named OCR-CC, which contains 1:4 million images with scene text. Pre-trained on this OCR-CC dataset, our approach outperforms the state of the art by large margins on multiple tasks, i.e., +8:3% accuracy on TextVQA, +8:6% accuracy on ST-VQA, and +10:2 CIDEr score on TextCaps.