Crowdsourced Corpus of Sentence Simplification with Core Vocabulary

Crowdsourced Corpus of Sentence Simplification with Core Vocabulary
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Akihiro Katsuta;Kazuhide Yamamoto
Akihiro Katsuta;Kazuhide Yamamoto
中科院分区:
其他
文献类型:
--
作者:
Akihiro Katsuta;Kazuhide Yamamoto

文献摘要

相似文献

我们提出了一个新的日本众包数据集,由更复杂的句子创建的简化句子。我们的简化标准包括简化句子中所有可重写的单词都是从2000个单词的核心词汇中提取的。我们的简化语料库是日语教科书和参考书中的复杂句子的集合,以及人类生成的简化句子,以及如何解释复杂句子的数据。该语料库包含15,000个句子,包括复杂和简单的版本。此外,我们还研究了每个注释器使用的简化操作的差异。目的是了解众包的复杂-简单并行语料库是否是通过机器学习进行自动简化的合适数据源。结果是,在构建数据集的注释者之间存在高度的一致性。因此,我们相信这个语料库是一个用于机器学习简化的高质量数据集。因此,我们计划在未来扩大简化语料库的规模。
We present a new Japanese crowdsourced data set of simplified sentences created from more complex ones. Our simplicity standard involves all rewritable words in the simplified sentences being drawn from a core vocabulary of 2,000 words. Our simplified corpus is a collection of complex sentences from Japanese textbooks and reference books together with simplified sentences generated by humans, paired with data on how the complex sentences were paraphrased. The corpus contains a total of 15,000 sentences, in both complex and simple versions. In addition, we investigate the differences in the simplification operations used by each annotator. The aim is to understand whether a crowdsourced complex-simple parallel corpus is an appropriate data source for automated simplification by machine learning. The results, that there was a high level of agreement between the annotators building the data set. So, we believe that this corpus is a good quality data set for machine learning for simplification. We therefore plan to expand the scale of the simplified corpus in the future.