Morfessor EM+Prune: Improved Subword Segmentation with Expectation Maximization and Pruning

Morfessor EM+Prune: Improved Subword Segmentation with Expectation Maximization and Pruning
复制标题

Morfessor EM Prune:通过期望最大化和修剪改进子词分割

DOI:
--
复制
发表时间:
2020
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
M. Kurimo
M. Kurimo
中科院分区:
--
文献类型:
--
作者:
Stig;Sami Virpioja;M. Kurimo

文献摘要

被引文献

相似文献

数据驱动的分词到子词单位的方法已经在自动语音识别和统计机器翻译等各种自然语言处理应用中使用了近20年。最近,它被广泛采用,因为基于深度神经网络的模型经常受益于子词单位,甚至对于形态学较简单的语言也是如此。在本文中,我们讨论并比较了基于期望最大化算法和词典修剪的一元子词模型的训练算法。通过使用英语、芬兰语、北萨米语和土耳其语数据集,我们发现这种方法能够找到更好的解决方案来解决由Morfessor Baseline模型定义的优化问题,而不是其原始的递归训练算法。与语言金标准相比,改进的优化还导致更高的形态学分割精度。我们在广泛使用的教授软件包中发布了新算法的实现。
Data-driven segmentation of words into subword units has been used in various natural language processing applications such as automatic speech recognition and statistical machine translation for almost 20 years. Recently it has became more widely adopted, as models based on deep neural networks often benefit from subword units even for morphologically simpler languages. In this paper, we discuss and compare training algorithms for a unigram subword model, based on the Expectation Maximization algorithm and lexicon pruning. Using English, Finnish, North Sami, and Turkish data sets, we show that this approach is able to find better solutions to the optimization problem defined by the Morfessor Baseline model than its original recursive training algorithm. The improved optimization also leads to higher morphological segmentation accuracy when compared to a linguistic gold standard. We publish implementations of the new algorithms in the widely-used Morfessor software package.