Normalisation of Historical Text Using Context-Sensitive Weighted Levenshtein Distance and Compound Splitting

Normalisation of Historical Text Using Context-Sensitive Weighted Levenshtein Distance and Compound Splitting
复制标题

使用上下文敏感加权编辑距离和复合分裂对历史文本进行标准化

DOI:
--
复制
发表时间:
2013
期刊:
Nordic Conference of Computational Linguistics
影响因子:
--
通讯作者:
Joakim Nivre
Joakim Nivre
中科院分区:
--
文献类型:
--
作者:
Eva Pettersson;Beáta Megyesi;Joakim Nivre

文献摘要

被引文献

相似文献

历史文本的自然语言处理带来了各种挑战,例如处理高度的拼写变化。此外,通常没有足够的语言注释数据可用于训练词性标记器和其他旨在处理这种特定类型文本的工具。在本文中,我们提出了一个Levenshtein为基础的方法来规范化的历史文本的现代拼写。这使我们能够将在当代语料库上训练的标准NLP工具应用于历史输入文本的规范化版本。在其基本版本中,不需要注释的历史数据,因为用于Levenshtein比较的唯一数据是当代词典或语料库。此外,可以可选地包括手动规范化的历史文本的(小)语料库,以便以监督的方式学习频繁单词的规范化和编辑操作的权重,这提高了精度。我们表明,这种方法是成功的规范化的准确性方面,并通过一个标准的现代标签应用到历史文本的性能。我们还比较了我们的方法,以前实现的方法,使用一组手写的规范化规则,我们看到,Levenshtein为基础的方法显然优于手工制作的规则。此外,瑞典数据进行了实验,结果令人鼓舞,我们相信我们的方法可以成功地适用于分析其他语言的历史文本,包括那些资源较少的语言。
Natural language processing for historical text imposes a variety of challenges, such as to deal with a high degree of spelling variation. Furthermore, there is often not enough linguistically annotated data available for training part-of-speech taggers and other tools aimed at handling this specific kind of text. In this paper we present a Levenshtein-based approach to normalisation of historical text to a modern spelling. This enables us to apply standard NLP tools trained on contemporary corpora on the normalised version of the historical input text. In its basic version, no annotated historical data is needed, since the only data used for the Levenshtein comparisons are a contemporary dictionary or corpus. In addition, a (small) corpus of manually normalised historical text can optionally be included to learn normalisation for frequent words and weights for edit operations in a supervised fashion, which improves precision. We show that this method is successful both in terms of normalisation accuracy, and by the performance of a standard modern tagger applied to the historical text. We also compare our method to a previously implemented approach using a set of hand-written normalisation rules, and we see that the Levenshtein-based approach clearly outperforms the hand-crafted rules. Furthermore, the experiments were carried out on Swedish data with promising results and we believe that our method could be successfully applicable to analyse historical text for other languages, including those with less resources.