The corpus of Basque simplified texts (CBST)

The corpus of Basque simplified texts (CBST)
复制标题

巴斯克语简化文本语料库 (CBST)

DOI:
10.1007/s10579-017-9407-6
复制
发表时间:
2017
影响因子:
2.7
通讯作者:
Arantza Díaz de Ilarraza
Arantza Díaz de Ilarraza
中科院分区:
计算机科学4区
文献类型:
--
作者:
Itziar Gonzalez;M. J. Aranzabe;Arantza Díaz de Ilarraza

文献摘要

被引文献

相似文献

在本文中,我们提出了巴斯克简化文本的语料库。该语料库收录了科普领域的227个原文句子,每个句子有两个简化版本。简化版本是按照不同的方法创建的:结构性的,由法院翻译考虑易于阅读的指导方针,直观的,由教师根据她的经验。本语料库的目的是对简化文本进行比较分析。为此,我们还提出了我们创建的注释语料库的注释方案。注释方案分为八个宏操作:删除、合并、拆分、转换、插入、重新排序、无操作和其他。这些宏操作可以分为不同的操作。我们还将我们的工作和成果与其他语言相关联。该语料库将用于证实所作的决定,并改进巴斯克语文本自动简化系统的设计。
In this paper we present the corpus of Basque simplified texts. This corpus compiles 227 original sentences of science popularisation domain and two simplified versions of each sentence. The simplified versions have been created following different approaches: the structural, by a court translator who considers easy-to-read guidelines and the intuitive, by a teacher based on her experience. The aim of this corpus is to make a comparative analysis of simplified text. To that end, we also present the annotation scheme we have created to annotate the corpus. The annotation scheme is divided into eight macro-operations: delete, merge, split, transformation, insert, reordering, no operation and other. These macro-operations can be classified into different operations. We also relate our work and results to other languages. This corpus will be used to corroborate the decisions taken and to improve the design of the automatic text simplification system for Basque.