Superbizarre Is Not Superb: Derivational Morphology Improves BERT’s Interpretation of Complex Words

Superbizarre Is Not Superb: Derivational Morphology Improves BERT’s Interpretation of Complex Words
复制标题

DOI:
10.18653/v1/2021.acl-long.279
复制
发表时间:
2021-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Valentin Hofmann;J. Pierrehumbert;Hinrich Schütze
Valentin Hofmann;J. Pierrehumbert;Hinrich Schütze
中科院分区:
其他
文献类型:
--
作者:
Valentin Hofmann;J. Pierrehumbert;Hinrich Schütze

文献摘要

相似文献

预训练语言模型 (PLM) 的输入分段如何影响其对复杂单词的解释?我们提出了第一个研究这个问题的研究,以 BERT 作为 PLM 的例子,重点关注其英语衍生词的语义表示。我们证明 PLM 可以解释为串行双路由模型,即复杂单词的含义要么被存储,要么需要从子词计算,这意味着最大有意义的输入标记应该允许对新单词进行最佳泛化。这一假设得到了一系列语义探测任务的证实,在这些任务上,DelBERT(利用 BERT 进行推导)(一种具有推导输入分割的模型)的性能大大优于带有 WordPiece 分割的 BERT。我们的结果表明,如果使用输入标记的形态信息词汇表,PLM 的泛化能力可以进一步提高。
How does the input segmentation of pretrained language models (PLMs) affect their interpretations of complex words? We present the first study investigating this question, taking BERT as the example PLM and focusing on its semantic representations of English derivatives. We show that PLMs can be interpreted as serial dual-route models, i.e., the meanings of complex words are either stored or else need to be computed from the subwords, which implies that maximally meaningful input tokens should allow for the best generalization on new words. This hypothesis is confirmed by a series of semantic probing tasks on which DelBERT (Derivation leveraging BERT), a model with derivational input segmentation, substantially outperforms BERT with WordPiece segmentation. Our results suggest that the generalization capabilities of PLMs could be further improved if a morphologically-informed vocabulary of input tokens were used.