Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer

Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
复制标题

DOI:
10.1007/978-3-030-86331-9_47
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Rafal Powalski;Łukasz Borchmann;Dawid Jurkiewicz;Tomasz Dwojak;Michal Pietruszka;Gabriela Pałka
Rafal Powalski;Łukasz Borchmann;Dawid Jurkiewicz;Tomasz Dwojak;Michal Pietruszka;Gabriela Pałka
中科院分区:
其他
文献类型:
--
作者:
Rafal Powalski;Łukasz Borchmann;Dawid Jurkiewicz;Tomasz Dwojak;Michal Pietruszka;Gabriela Pałka

文献摘要

被引文献

相似文献

我们通过引入同时学习布局信息、视觉特征和文本语义的TILT神经网络架构,解决了超越纯文本文档的自然语言理解的挑战性问题。与以前的方法相反,我们依赖于能够统一涉及自然语言的各种问题的解码器。该布局被表示为注意偏差,并辅以上下文化的视觉信息,而我们模型的核心是一个预训练的编码器-解码器转换器。我们的新方法在从文档中提取信息和回答需要布局理解的问题(DocVQA, CORD, SROIE)方面取得了最先进的结果。同时,我们通过采用端到端模型简化了流程。
We address the challenging problem of Natural Language Comprehension beyond plain-text documents by introducing the TILT neural network architecture which simultaneously learns layout information, visual features, and textual semantics. Contrary to previous approaches, we rely on a decoder capable of unifying a variety of problems involving natural language. The layout is represented as an attention bias and complemented with contextualized visual information, while the core of our model is a pretrained encoder-decoder Transformer. Our novel approach achieves state-of-the-art results in extracting information from documents and answering questions which demand layout understanding (DocVQA, CORD, SROIE). At the same time, we simplify the process by employing an end-to-end model.