Attending to Long-Distance Document Context for Sequence Labeling

Attending to Long-Distance Document Context for Sequence Labeling
复制标题

DOI:
10.18653/v1/2020.findings-emnlp.330
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Matthew Jörke;Jon Gillick;Matthew Sims;David Bamman
Matthew Jörke;Jon Gillick;Matthew Sims;David Bamman
中科院分区:
其他
文献类型:
--
作者:
Matthew Jörke;Jon Gillick;Matthew Sims;David Bamman

文献摘要

相似文献

在这项工作中,我们提出了一种方法,将全球范围内的长文档时,在序列标签问题,如NER的本地决策。受特征化对数线性模型工作的启发(Chieu and Ng,2002;萨顿and McCallum,2004),我们的模型学习在为上下文中的每个标记生成表示时关注相同单词类型的多次提及,将该工作扩展到学习可以纳入现代神经模型的表示。在测试时关注更广泛的背景为预训练提供了补充信息(Gururangan等人,2020),比缺乏这种背景的等效参数化模型产生更强的增益,并且在识别具有高TF-IDF分数的实体(即,在一个文件中是重要的)。
We present in this work a method for incorporating global context in long documents when making local decisions in sequence labeling problems like NER. Inspired by work in featurized log-linear models (Chieu and Ng, 2002; Sutton and McCallum, 2004), our model learns to attend to multiple mentions of the same word type in generating a representation for each token in context, extending that work to learning representations that can be incorporated into modern neural models. Attending to broader context at test time provides complementary information to pretraining (Gururangan et al., 2020), yields strong gains over equivalently parameterized models lacking such context, and performs best at recognizing entities with high TF-IDF scores (i.e., those that are important within a document).