Using the Levenshtein algorithm for automatic lemmatization in Old English

Using the Levenshtein algorithm for automatic lemmatization in Old English
复制标题

使用 Levenshtein 算法进行古英语自动词形还原

DOI:
--
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
Bernadette Johnson
Bernadette Johnson
中科院分区:
--
文献类型:
--
作者:
Bernadette Johnson

文献摘要

被引文献

相似文献

本研究旨在开发和测试一个自动词形还原程序,用于古英语语言,利用Levenshtein编辑距离算法,词干提取和其他技术来帮助克服拼写不规则和屈折词尾等问题。主要目标是创建一个词元列表,导入文本分析软件,将古英语单词与其变体等同起来,这一目标取得了有限但有希望的成功。主要的词形化程序是用Perl编程语言编写的。Perl和Perl中的其他脚本,以及Unix命令行命令和AntConc 3.2.0语料库分析软件,用于从古英语语料库XML文件的字典中提取文本文件,生成可用文本中所有单词的排序列表,供主程序使用,并操作和分析数据。
This study was undertaken to develop and test an automatic lemmatization program for the Old English language utilizing the Levenshtein edit distance algorithm, stemming, and other techniques to help overcome issues such as rampant spelling irregularity and the presence of inflectional endings. The primary goal was to create a lemma list for import into text analysis software to equate Old English words with their variants, which was met with limited but promising success. The main lemmatization program is written in the Perl programming language. Other scripts in Perl and XSLT, as well as Unix command line commands and AntConc 3.2.0 corpus analysis software, were used to extract text files from the Dictionary of Old English Corpus XML files, to generate sorted lists of all the words in the available texts for use by the main program, and to manipulate and analyze the data.