BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents
BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents
复制标题
BROS:一种预训练的语言模型,专注于文本和布局,以便更好地从文档中提取关键信息
DOI:
10.1609/aaai.v36i10.21322
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Sungrae Park
中科院分区:
文献类型:
--
作者:
Teakgyu Hong;Donghyun Kim;Mingi Ji;Wonseok Hwang;Daehyun Nam;Sungrae Park
Key information extraction (KIE) from document images requires understanding the contextual and spatial semantics of texts in two-dimensional (2D) space.
Many recent studies try to solve the task by developing pre-trained language models focusing on combining visual features from document images with texts and their layout.
On the other hand, this paper tackles the problem by going back to the basic: effective combination of text and layout.
Specifically, we propose a pre-trained language model, named BROS (BERT Relying On Spatiality), that encodes relative positions of texts in 2D space and learns from unlabeled documents with area-masking strategy.
With this optimized training scheme for understanding texts in 2D space, BROS shows comparable or better performance compared to previous methods on four KIE benchmarks (FUNSD, SROIE*, CORD, and SciTSR) without relying on visual features.
This paper also reveals two real-world challenges in KIE tasks--(1) minimizing the error from incorrect text ordering and (2) efficient learning from fewer downstream examples--and demonstrates the superiority of BROS over previous methods.
DOI:
10.1007/978-3-030-86331-9_47
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
作者:
Rafal Powalski;Łukasz Borchmann;Dawid Jurkiewicz;Tomasz Dwojak;Michal Pietruszka;Gabriela Pałka
通讯作者:
Rafal Powalski;Łukasz Borchmann;Dawid Jurkiewicz;Tomasz Dwojak;Michal Pietruszka;Gabriela Pałka