SpanBERT: Improving Pre-training by Representing and Predicting Spans

SpanBERT: Improving Pre-training by Representing and Predicting Spans
复制标题

DOI:
10.1162/tacl_a_00300
复制
发表时间:
2020-01-01
影响因子:
10.9
通讯作者:
Levy, Omer
Levy, Omer
中科院分区:
人文科学1区
文献类型:
--
作者:
Joshi, Mandar;Chen, Danqi;Levy, Omer

文献摘要

被引文献

相似文献

我们提出了SpanBERT,这是一种预训练方法,旨在更好地表示和预测文本的跨度。我们的方法通过以下方式扩展了BERT:(1)掩蔽连续的随机跨度,而不是随机令牌,(2)训练跨度边界表示来预测掩蔽跨度的整个内容,而不依赖于其中的单个令牌表示。SpanBERT始终优于BERT和我们更好的调整基线,在回答问题和共指消解等跨度选择任务上有很大的收益。特别是,在与BERTlarge相同的训练数据和模型大小的情况下,我们的单个模型在SQuAD 1.1和2.0上分别获得94.6%和88.7%的F1。我们还实现了OntoNotes共指消解任务(79.6%F1)的最新技术水平,在TACRED关系提取基准测试中表现出色,甚至在GLUE上也有收获。(一)
We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the span boundary representations to predict the entire content of the masked span, without relying on the individual token representations within it. SpanBERT consistently outperforms BERT and our better-tuned baselines, with substantial gains on span selection tasks such as question answering and coreference resolution. In particular, with the same training data and model size as BERTlarge, our single model obtains 94.6% and 88.7% F1 on SQuAD 1.1 and 2.0 respectively. We also achieve a new state of the art on the OntoNotes coreference resolution task (79.6% F1), strong performance on the TACRED relation extraction benchmark, and even gains on GLUE.(1)