Using Semantics for Granularities of Tokenization

Using Semantics for Granularities of Tokenization
复制标题

DOI:
10.1162/coli_a_00325
复制
发表时间:
2018-09-01
影响因子:
9.3
通讯作者:
Biemann, Chris
Biemann, Chris
中科院分区:
计算机科学3区
文献类型:
--
作者:
Riedl, Martin;Biemann, Chris

文献摘要

被引文献

相似文献

根据下游应用,建议将标记化的概念从基于低级字符的标记边界检测扩展到识别有意义和有用的语言单元。这既需要识别由几个组成a的单个单词组成的单位,也需要将单个单词复合词拆分成它们有意义的部分。在本文中,我们介绍了这两个任务的无监督方法和无知识方法。我们的研究的主要新颖性是基于这样一个事实,即方法主要基于分布相似度,其中我们使用了两种风格:基于稀疏计数的分布语义模型和基于密集神经的分布语义模型。首先,我们介绍了一种检测MWES的方法--Druid。在两种语言的MWE标注数据集和新提取的32种语言的评价数据集上的评估结果表明,Druid比以前的方法更好地利用了分布信息。其次,我们提出了一种分解封闭化合物的算法--SECOS。在对四种语言的四个专用分解数据集以及从维基词典中提取的14种语言的数据集的评估中,我们展示了我们的方法相对于非监督基线的优越性,有时甚至与以前的特定语言和监督方法的性能相当。在最后的实验中,我们展示了如何将分解和MWE信息用于信息检索。在这里,当我们将单词信息与MWES和复合部分结合在一起时,我们获得了最好的结果。总体而言,我们的方法为自动检测词汇单元铺平了道路,而不是标准的标记化技术,而不需要特定语言的预处理步骤,如词性标记。
Depending on downstream applications, it is advisable to extend the notion of tokenization from low-level character-based token boundary detection to identification of meaningful and useful language units. This entails both identifying units composed of several single words that form a several single words that form a, as well as splitting single-word compounds into their meaningful parts. In this article, we introduce unsupervised and knowledge-free methods for these two tasks. The main novelty of our research is based on the fact that methods are primarily based on distributional similarity, of which we use two flavors: a sparse count-based and a dense neural-based distributional semantic model. First, we introduce DRUID, which is a method for detecting MWEs. The evaluation on MWE-annotated data sets in two languages and newly extracted evaluation data sets for 32 languages shows that DRUID compares favorably over previous methods not utilizing distributional information. Second, we present SECOS, an algorithm for decompounding close compounds. In an evaluation of four dedicated decompounding data sets across four languages and on data sets extracted from Wiktionary for 14 languages, we demonstrate the superiority of our approach over unsupervised baselines, sometimes even matching the performance of previous language-specific and supervised methods. In a final experiment, we show how both decompounding and MWE information can be used in information retrieval. Here, we obtain the best results when combining word information with MWEs and the compound parts in a bag-of-words retrieval set-up. Overall, our methodology paves the way to automatic detection of lexical units beyond standard tokenization techniques without language-specific preprocessing steps such as POS tagging.