Normalising Slovene data: historical texts vs. user-generated content

Normalising Slovene data: historical texts vs. user-generated content
复制标题

标准化斯洛文尼亚数据:历史文本与用户生成的内容

DOI:
--
复制
发表时间:
2016
期刊:
Conference on Natural Language Processing
影响因子:
--
通讯作者:
T. Erjavec
T. Erjavec
中科院分区:
--
文献类型:
--
作者:
Nikola Ljubesic;Katja Zupan;Darja Fišer;T. Erjavec

文献摘要

被引文献

相似文献

本文介绍了两个手动注释的斯洛文尼亚语文本规范化数据集,一个是历史文本,另一个是推文,并提出了几种基于字符的统计机器翻译变体,以规范其单词的拼写。这些系统的不同之处在于它们是执行标记级规范化还是段级规范化,以及它们是否使用额外的语言资源。这些系统将根据黄金标准进行自动评估,并根据新开发的错误类型进行人工评估,以详细分析不同类型数据和不同数据标准化水平的影响。评估表明,段级归一化可以是有用的给定足够高的水平的token歧义,可以使用相同的系统,无论数据类型,和背景资源将始终被证明是有用的。
The paper presents two manually annotated Slovene language text normalisation datasets, one of historical texts and the other of tweets, and proposes several variants of character-based statistical machine translation to normalise the spelling of their words. The systems differ in whether they perform token-level or segment-level normalisation and whether they make use of additional language resources. The systems are evaluated automatically against the gold standard as well as manually, against a newly developed typology of errors intended to analyse in detail the effect of different types of data and different levels of data standardness. The evaluations show that segment-level normalisation can be useful given a high enough level of to-ken ambiguity, that the same system can be used regardless of the data type, and that background resources will always prove useful.