Old English class I strong verbs lemmatisation: A morphological generation approach

Old English class I strong verbs lemmatisation: A morphological generation approach
复制标题

古英语 I 类强动词词形还原:形态生成方法

DOI:
10.1080/00393274.2021.2010128
复制
发表时间:
2022
影响因子:
0.4
通讯作者:
Roberto Torre Alonso
Roberto Torre Alonso
中科院分区:
--
文献类型:
--
作者:
Roberto Torre Alonso

文献摘要

被引文献

相似文献

摘要 本文对古英语 I 类强动词 L-Y 的基于类型的自动词形还原提出了问题。为了实现这一目标,本文讨论了形态生成界面的设计,该界面使用古英语强动词变形规则以及最常见的变体拼写来实现。界面提供的表格会自动与两个最具代表性的古英语语料库(古英语词典语料库和约克-多伦多-赫尔辛基古英语散文解析语料库)进行检查,并将验证结果分配给相应的引理。结果证明,几乎 99% 的经过验证的形式都可以成功分配引理。剩下的 1% 对应于形式上不明确的屈折形式,其中引理分配在两个候选人之间存在争议。得出的结论是,只有在基于上下文的标记分析的基础上才能消除歧义。该文章证实,古英语强动词基于类型的词形还原可以在很大程度上实现自动化。
ABSTRACT This article takes issue with type-based automatic lemmatisation of the Old English class I strong verbs L-Y. To reach this goal, the article discusses the design of a Morphological Generation interface implemented with the rules of Old English strong verb inflection as well as with the most frequent variant spellings. The forms provided by the interface are automatically checked against the two most representative corpora of Old English, the Dictionary of Old English Corpus and the The York-Toronto-Helsinki Parsed Corpus of Old English Prose, and validated results are assigned the corresponding lemma. The results prove that almost 99% of the validated forms can successfully be assigned lemma. The remaining 1% correspond to formally ambiguous inflectional forms, where lemma assignment is in dispute between two candidates. The conclusion is reached that disambiguation is only possible on the basis of contextualised token-based analysis. The article confirms that type-based lemmatisation of Old English strong verbs can be largely automatised.