Morphological Lexicon Extraction from Raw Text Data

Morphological Lexicon Extraction from Raw Text Data
复制标题

从原始文本数据中提取形态词典

DOI:
10.1007/11816508_49
复制
发表时间:
2006
期刊:
FinTAL
影响因子:
--
通讯作者:
Aarne Ranta
Aarne Ranta
中科院分区:
--
文献类型:
--
作者:
Markus Forsberg;H. Hammarström;Aarne Ranta

文献摘要

被引文献

相似文献

该工具可从原始文本数据中自动提取引理-范式对。该工具使用由正则表达式和命题逻辑组成的搜索模式。这些搜索模式根据数据中出现的单词形式,定义了在词典中包括引理-范式对的充分条件。本文阐述了抽取词的搜索模式语法和搜索算法,并从查全率和查准率的角度讨论了搜索模式的设计。提取工具是为功能形态学工具[1]中定义的形态学开发的,但它可用于实现形态学的词-范式描述的所有系统。该工具的有用性通过对加拿大汉萨法语语料库的案例研究来证明。结果是根据提取引理的精度和统计覆盖率和规则生产力来评估的。竞争性抽取数据表明,在定制工具中使用人工编写的规则是解决手头任务的一种省时方法。
The toolextractenables the automatic extraction of lemma-paradigm pairs from raw text data. The tool uses search patterns that consist of regular expressions and propositional logic. These search patterns define sufficient conditions for including lemma-paradigm pairs in the lexicon, on the basis of word forms occurring in the data. This paper explains the search pattern syntax ofextractas well as the search algorithm, and discusses the design of search patterns from the recall and precision point of view.Theextracttool was developed for morphologies defined in theFunctional Morphologytool [1], but it is usable for all systems that implement a word-and-paradigm description of a morphology.The usefulness of the tool is demonstrated by a case study on the Canadian Hansards Corpus of French. The result is evaluated in terms of precision of the extracted lemmas and statistics on coverage and rule productiveness. Competitive extraction figures show that human-written rules in a tailored tool is a time-efficient approach to the task at hand.