Normalising Slovene data: historical texts vs. user-generated content
Normalising Slovene data: historical texts vs. user-generated content
复制标题
标准化斯洛文尼亚数据:历史文本与用户生成的内容
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
T. Erjavec
中科院分区:
文献类型:
--
作者:
Nikola Ljubesic;Katja Zupan;Darja Fišer;T. Erjavec
The paper presents two manually annotated Slovene language text normalisation datasets, one of historical texts and the other of tweets, and proposes several variants of character-based statistical machine translation to normalise the spelling of their words. The systems differ in whether they perform token-level or segment-level normalisation and whether they make use of additional language resources. The systems are evaluated automatically against the gold standard as well as manually, against a newly developed typology of errors intended to analyse in detail the effect of different types of data and different levels of data standardness. The evaluations show that segment-level normalisation can be useful given a high enough level of to-ken ambiguity, that the same system can be used regardless of the data type, and that background resources will always prove useful.