A Large-Scale Comparison of Historical Text Normalization Systems

A Large-Scale Comparison of Historical Text Normalization Systems
复制标题

历史文本规范化系统的大规模比较

DOI:
--
复制
发表时间:
2019
期刊:
North American Chapter of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Marcel Bollmann
Marcel Bollmann
中科院分区:
--
文献类型:
--
作者:
Marcel Bollmann

文献摘要

被引文献

相似文献

对于历史文本规范化的最新方法,目前还没有达成共识。已经提出了许多技术,包括基于规则的方法、距离度量、基于字符的统计机器翻译和神经编解码器模型,但研究使用了不同的数据集、不同的评估方法,并得出了不同的结论。本文是迄今为止规模最大的一次历史文本规范化研究。我们批判性地调查了现有的八种语言的文献和报告实验,比较了跨越所有类别的所提出的归一化技术的系统,分析了训练数据量的效果,并使用了不同的评估方法。数据集和脚本公开可用。
There is no consensus on the state-of-the-art approach to historical text normalization. Many techniques have been proposed, including rule-based methods, distance metrics, character-based statistical machine translation, and neural encoder–decoder models, but studies have used different datasets, different evaluation methods, and have come to different conclusions. This paper presents the largest study of historical text normalization done so far. We critically survey the existing literature and report experiments on eight languages, comparing systems spanning all categories of proposed normalization techniques, analysing the effect of training data quantity, and using different evaluation methods. The datasets and scripts are made publicly available.