Representation Learning for Information Extraction from Form-like Documents
Representation Learning for Information Extraction from Form-like Documents
复制标题
从类似表单的文档中提取信息的表示学习
DOI:
10.18653/v1/2020.acl-main.580
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Marc Najork
中科院分区:
文献类型:
--
作者:
Bodhisattwa Prasad Majumder;Navneet Potti;Sandeep Tata;James Bradley Wendt;Qi Zhao;Marc Najork
We propose a novel approach using representation learning for tackling the problem of extracting structured information from form-like document images. We propose an extraction system that uses knowledge of the types of the target fields to generate extraction candidates and a neural network architecture that learns a dense representation of each candidate based on neighboring words in the document. These learned representations are not only useful in solving the extraction task for unseen document templates from two different domains but are also interpretable, as we show using loss cases.
DOI:
10.1145/775152.775155
发表时间:
2003-05
期刊:
--
影响因子:
--
作者:
Shipeng Yu;Deng Cai;Ji-Rong Wen;Wei-Ying Ma
通讯作者:
Shipeng Yu;Deng Cai;Ji-Rong Wen;Wei-Ying Ma