Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Sabrina J. Mielke;Zaid Alyafeai;Elizabeth Salesky;Colin Raffel;Manan Dey;Matthias Gallé;Arun Raja;Chenglei Si;Wilson Y. Lee;Benoît Sagot;Samson Tan
Sabrina J. Mielke;Zaid Alyafeai;Elizabeth Salesky;Colin Raffel;Manan Dey;Matthias Gallé;Arun Raja;Chenglei Si;Wilson Y. Lee;Benoît Sagot;Samson Tan
中科院分区:
其他
文献类型:
--
作者:
Sabrina J. Mielke;Zaid Alyafeai;Elizabeth Salesky;Colin Raffel;Manan Dey;Matthias Gallé;Arun Raja;Chenglei Si;Wilson Y. Lee;Benoît Sagot;Samson Tan

文献摘要

相似文献

我们想要建模的文本单位是什么?从字节到多字表达式,可以以多种粒度分析和生成文本。直到最近,大多数自然语言处理(NLP)模型都是在单词上操作的,将这些单词视为离散的和原子的标记,但从字节对编码(BPE)开始,基于子词的方法已经在许多领域占据主导地位,支持较小的词汇量,同时仍然允许快速推理。道路的末尾是字符级模型还是字节级处理?在这项调查中,我们通过展示单词和字符的混合方法以及基于学习切分的基于子词的方法是如何被提出和评估的,将前神经时代和神经时代的几个工作联系在一起。我们的结论是,有可能永远不会有适用于所有应用程序的灵丹妙药,认真考虑令牌化对许多应用程序仍然很重要。
What are the units of text that we want to model? From bytes to multi-word expressions, text can be analyzed and generated at many granularities. Until recently, most natural language processing (NLP) models operated over words, treating those as discrete and atomic tokens, but starting with byte-pair encoding (BPE), subword-based approaches have become dominant in many areas, enabling small vocabularies while still allowing for fast inference. Is the end of the road character-level model or byte-level processing? In this survey, we connect several lines of work from the pre-neural and neural era, by showing how hybrid approaches of words and characters as well as subword-based approaches based on learned segmentation have been proposed and evaluated. We conclude that there is and likely will never be a silver bullet singular solution for all applications and that thinking seriously about tokenization remains important for many applications.