A Two-level Morphological Analyser and Generator for Irish using Finite-State Transducers

A Two-level Morphological Analyser and Generator for Irish using Finite-State Transducers
复制标题

使用有限状态换能器的爱尔兰语两级形态分析器和生成器

DOI:
--
复制
发表时间:
2002
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Elaine Uí Dhonnchadha
Elaine Uí Dhonnchadha
中科院分区:
--
文献类型:
--
作者:
Elaine Uí Dhonnchadha

文献摘要

被引文献

相似文献

计算形态学是自然语言处理的重要组成部分。有限状态技术已成功应用于世界上许多主要语言的计算音系学和形态学。现代爱尔兰语等凯尔特语言呈现出具有挑战性的形态特征,迄今为止尚未使用有限状态技术来解决这些特征。本文提出了使用施乐有限状态工具开发的爱尔兰语的有限状态两级形态学。该系统对现代爱尔兰语中所有变形词性的变形形态进行编码。词干和词缀的词法被编码在词典中,并且单词突变被实现为一系列编码为正则表达式的替换规则。词典和规则都被编译成有限状态转换器,并组合起来生成该语言的单个词汇转换器。形态学的有限状态两级实现的一个主要优点是它们固有的双向性;同一系统用于分析和生成该语言的词形。该资源可用作许多 NLP 应用程序的组成部分,例如拼写检查器/校正器、词干分析器和文本到语音合成器。它还可以用于文本语料库的标记化、词形还原和词性标记。该系统旨在广泛覆盖该语言,并根据当代爱尔兰文本语料库中最常用的单词进行评估。最后,建议对该系统进行可能的扩展,例如派生形态和包含方言或历史词形。
Computational morphology is an important part of natural language processing. Finite-state techniques have been applied successfully in computational phonology and morphology to many of the world’s major languages. Celtic languages such as Modern Irish present challenging morphological features that to date have not been addressed using finite-state technology. This paper presents a finite-state two-level morphology of Irish developed using Xerox Finite-State Tools. The system encodes the inflectional morphology of all inflected parts-of-speech in Modern Irish. The morphotactics of stems and affixes are encoded in the lexicon and word mutations are implemented as a series of replace rules encoded as regular expressions. Both the lexicons and rules are compiled into finite state transducers and combined to produce a single lexical transducer for the language. A major advantage of finite-state two-level implementations of morphology is their inherent bi-directionality; the same system is used for both analysis and generation of word forms in the language. This resource can be used as a component part in many NLP applications such as spelling checkers/correctors, stemmers, and text to speech synthesisers. It can also be used in tokenising, lemmatising and part-of-speech tagging of a corpus of text. The system, which is designed for broad coverage of the language, is evaluated against the most frequently used words in a corpus of contemporary Irish texts. Finally, possible extensions to the system are suggested, such as derivational morphology and the inclusion of dialectal or historical word-forms.