Normalisation of Historical Text Using Context-Sensitive Weighted Levenshtein Distance and Compound Splitting
Normalisation of Historical Text Using Context-Sensitive Weighted Levenshtein Distance and Compound Splitting
复制标题
使用上下文敏感加权编辑距离和复合分裂对历史文本进行标准化
DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Joakim Nivre
中科院分区:
文献类型:
--
作者:
Eva Pettersson;Beáta Megyesi;Joakim Nivre
Natural language processing for historical text imposes a variety of challenges, such as to deal with a high degree of spelling variation. Furthermore, there is often not enough linguistically annotated data available for training part-of-speech taggers and other tools aimed at handling this specific kind of text. In this paper we present a Levenshtein-based approach to normalisation of historical text to a modern spelling. This enables us to apply standard NLP tools trained on contemporary corpora on the normalised version of the historical input text. In its basic version, no annotated historical data is needed, since the only data used for the Levenshtein comparisons are a contemporary dictionary or corpus. In addition, a (small) corpus of manually normalised historical text can optionally be included to learn normalisation for frequent words and weights for edit operations in a supervised fashion, which improves precision. We show that this method is successful both in terms of normalisation accuracy, and by the performance of a standard modern tagger applied to the historical text. We also compare our method to a previously implemented approach using a set of hand-written normalisation rules, and we see that the Levenshtein-based approach clearly outperforms the hand-crafted rules. Furthermore, the experiments were carried out on Swedish data with promising results and we believe that our method could be successfully applicable to analyse historical text for other languages, including those with less resources.