Using the Levenshtein Edit Distance for Automatic Lemmatization: A Case Study for Modern Greek and English

Using the Levenshtein Edit Distance for Automatic Lemmatization: A Case Study for Modern Greek and English
复制标题

使用编辑编辑距离进行自动词形还原:现代希腊语和英语的案例研究

DOI:
--
复制
发表时间:
2007
期刊:
IEEE International Conference on Tools with Artificial Intelligence
影响因子:
--
通讯作者:
N. Fakotakis
N. Fakotakis
中科院分区:
--
文献类型:
--
作者:
Dimitrios P. Lyras;K. Sgarbas;N. Fakotakis

文献摘要

被引文献

相似文献

在本工作中,我们在一个基于词典的算法上实现了编辑距离(也称为Levenshtein距离),以便在没有直接监督的情况下实现对规则和轻微不规则词的规格化形式(引理)的自动归纳。该算法结合了基于字符串相似度和最频繁屈折后缀的两种对齐模型。在我们的实验中,我们还通过评估该算法在现代希腊语和英语语言上的性能,检验了该算法的语言无关性(即特定语法和语言规则的无关性)。结果非常有希望,因为我们实现了希腊语95%以上的准确率和英语96%以上的准确率。该算法可用于各种文本挖掘和语言应用,如拼写检查器、电子词典、词法分析器、搜索引擎等。
In the present work we have implemented the Edit Distance (also known as Levenshtein Distance) on a dictionary-based algorithm in order to achieve the automatic induction of the normalized form (lemma) of regular and mildly irregular words with no direct supervision. The algorithm combines two alignment models based on the string similarity and the most frequent inflexional suffixes. In our experiments, we have also examined the language-independency (i.e. independency of the specific grammar and inflexional rules of the language) of the presented algorithm by evaluating its performance on the Modern Greek and English languages. The results were very promising as we achieved more than 95 % of accuracy for the Greek language and more than 96 % for the English language. This algorithm may be useful to various text mining and linguistic applications such as spell-checkers, electronic dictionaries, morphological analyzers, search engines etc.