Weigh your words - memory-based lemmatization for Middle Dutch

Weigh your words - memory-based lemmatization for Middle Dutch
复制标题

权衡你的用词 - 基于记忆的中古荷兰语词形还原

DOI:
--
复制
发表时间:
2010
期刊:
Literary and Linguistic Computing
影响因子:
--
通讯作者:
G. Pauw
G. Pauw
中科院分区:
--
文献类型:
--
作者:
M. Kestemont;Walter Daelemans;G. Pauw

文献摘要

被引文献

相似文献

本文论述的是中古荷兰文学的词形化。这个文本集和其他中世纪的语料库一样,其特点是拼写差异巨大,这使得对这类数据进行计算分析变得困难。因此,词形化是许多应用程序中必不可少的预处理步骤,因为它允许从表面的文本变化中抽象出来,例如在拼写中。我们将使用的数据是Corpus-Gysseling,其中包含公元1300年之前所有幸存的中世纪荷兰文学手稿。在这篇文章中,我们将提出一个独立于语言的系统,可以“学习”内词元拼写变化。我们描述了一系列的实验与这个系统,使用基于内存的机器学习,并提出了两个解决方案,我们的数据的词形化:第一个过程试图生成新的拼写变体,第二个试图实现一个新的字符串距离度量,以更好地检测拼写变体。后一种系统试图通过经典的Levenshtein距离对候选人进行重新排序,从而大大提高了词形化的准确性。这一研究结果是令人鼓舞的,意味着在计算研究中的中古荷兰文学迈出了实质性的一步。我们的技术可能会感兴趣的其他研究领域,以及由于其语言独立的性质。1中古荷兰语的拼写变异中古荷兰语是一种历史语言的典型例子,显示出相当数量的拼写变异(货车der Voort货车der Kleij,2005; Ernst-Gerlach and Fuhr,2006; Kestemont and货车Dalen-Oskam,2009; Souvay and Pierrel,2009)。特别是在印刷机出现之前,荷兰语没有标准的语言变体,更不用说标准的拼写了。因此,中世纪荷兰语的拼写通常是高度语音和“个人”的性质,因为它会代表每个作家自己的方言发音和当地的拼写习惯。这就是为什么即使是非常频繁的单词也可以用非常不同的方式拼写,反映了低地国家的方言和地方标准的丰富多样性(图1)。这种拼写变化使得在任何计算应用程序中处理中世纪文本变得困难。例如,对于作者归属,其通信:Mike Kestemont,Universiteit安特卫彭,Stadscampus,Prinsstraat 13,Room D.118,2000安特卫彭,比利时。E-mail:mike. ua.ac.be Literary and Linguistic Computing,Vol. 25,No. 3,2010. !作者2010。由牛津大学出版社代表ALLC和ACH出版。All rights reserved.如需查阅,请发送电子邮件至:journals. oxfordjournals.org 287 doi:10.1093/llc/fqq 011 Advance Access published on 4 August 2010 at U nivrsiteit Antw epen Bibotheek on August 0,2010 llc.oxfjournals.org D ow naded rom
This article deals with the lemmatization of Middle Dutch literature. This text collection—like any other medieval corpus—is characterized by an enormous spelling variation, which makes it difficult to perform a computational analysis of this kind of data. Lemmatization is therefore an essential preprocessing step in many applications, since it allows the abstraction from superficial textual variation, for instance in spelling. The data we will work with is the Corpus-Gysseling, containing all surviving Middle Dutch literary manuscripts dated before 1300 AD. In this article we shall present a language-independent system that can ‘learn’ intra-lemma spelling variation. We describe a series of experiments with this system, using Memory-Based Machine Learning and propose two solutions for the lemmatization of our data: the first procedure attempts to generate new spelling variants, the second one seeks to implement a novel string distance metric to better detect spelling variants. The latter system attempts to rerank candidates suggested by a classic Levenshtein distance, leading to a substantial gain in lemmatization accuracy. This research result is encouraging and means a substantial step forward in the computational study of Middle Dutch literature. Our techniques might be of interest to other research domains as well because of their language-independent nature. 1 Spelling Variation in Middle Dutch Middle Dutch is a typical example of a historical language displaying a considerable amount of spelling variation (Van der Voort van der Kleij, 2005; Ernst-Gerlach and Fuhr, 2006; Kestemont and Van Dalen-Oskam, 2009; Souvay and Pierrel, 2009). Especially before the advent of the printing press, there existed no standard language variety of Dutch, let alone a standard spelling. As such, medieval Dutch spelling was generally highly phonological and ‘personal’ in nature, since it would represent each writer’s own dialectal pronunciation and local spelling habits. That is why even highly frequent words could be spelled in very different ways, reflecting the abundant variety of dialects and local substandards then found in the Low Countries (Fig. 1). This spelling variation makes it difficult to process medieval texts in any computational application. For instance for authorship attribution, it Correspondence: Mike Kestemont, Universiteit Antwerpen, Stadscampus, Prinsstraat 13, Room D.118, 2000 Antwerpen, Belgium. E-mail: mike.kestemont@ua.ac.be Literary and Linguistic Computing, Vol. 25, No. 3, 2010. ! The Author 2010. Published by Oxford University Press on behalf of ALLC and ACH. All rights reserved. For Permissions, please email: journals.permissions@oxfordjournals.org 287 doi:10.1093/llc/fqq011 Advance Access published on 4 August 2010 at U nivrsiteit Antw epen Bibotheek on Agust 0, 2010 llc.oxfjournals.org D ow naded rom