Manual and semi-automatic normalization of historical spelling - case studies from Early New High German

Manual and semi-automatic normalization of historical spelling - case studies from Early New High German
复制标题

历史拼写的手动和半自动标准化 - 早期新高地德语的案例研究

DOI:
--
复制
发表时间:
2012
期刊:
Conference on Natural Language Processing
影响因子:
--
通讯作者:
Florian Petran
Florian Petran
中科院分区:
--
文献类型:
--
作者:
Marcel Bollmann;Stefanie Dipper;J. Krasselt;Florian Petran

文献摘要

被引文献

相似文献

本文介绍了人工和半自动规范化的历史语言数据的工作。我们首先讨论我们用于映射历史到现代单词形式的指导方针。这些准则区分了规范化(倾向于接近原文的形式)和现代化(倾向于接近现代语言的形式)。平均注释者之间的协议是88.38%的一组数据从早期新高地德语。然后,我们介绍了Norma,一个半自动的归一化工具。它集成了不同的模块(词典查找,重写规则),以交互方式规范化单词。该工具在给定新输入的情况下动态更新规则条目集。根据文本和训练设置,标准化1,000个令牌的总体准确率为61.78 - 79.65%(基线:24.76 - 59.53%)。
This paper presents work on manual and semi-automatic normalization of historical language data. We first address the guidelines that we use for mapping historical to modern word forms. The guidelines distinguish between normalization (preferring forms close to the original) and modernization (preferring forms close to modern language). Average inter-annotator agreement is 88.38% on a set of data from Early New High German. We then present Norma, a semi-automatic normalization tool. It integrates different modules (lexicon lookup, rewrite rules) for normalizing words in an interactive way. The tool dynamically updates the set of rule entries, given new input. Depending on the text and training settings, normalizing 1,000 tokens results in overall accuracies of 61.78‐79.65% (baseline: 24.76‐59.53%).