Proposal of Japanese Vocabulary Difficulty Level Dictionaries for Automated Essay Scoring Support System Using Rubric

Proposal of Japanese Vocabulary Difficulty Level Dictionaries for Automated Essay Scoring Support System Using Rubric
复制标题

DOI:
10.1007/s40305-019-00270-z
复制
发表时间:
2019-11
影响因子:
1.4
通讯作者:
Megumi Yamamoto;Nobuo Umemura;H. Kawano
Megumi Yamamoto;Nobuo Umemura;H. Kawano
中科院分区:
数学4区
文献类型:
--
作者:
Megumi Yamamoto;Nobuo Umemura;H. Kawano

文献摘要

相似文献

我们正在开发一个Moodle插件,这是一个用于大学生基础教育的AES(自动作文评分)支持系统。我们的论文评价体系以评分标准为基础,分为内容、结构、证据、风格、技巧五个评价点。词汇水平是技能得分项目之一。它是使用Sunakawa等人构建的日语学习词典来计算的。由于这并没有完全涵盖学生水平的作文中使用的词汇,我们发现词汇水平评分的准确性存在问题。本文提出以日语维基百科为语料库,构建日语词汇难度水平综合词典。我们将潜在狄利克雷分配(LDA)应用于维基百科语料库,发现单词出现概率是衡量单词难度的指标之一。我们使用Tf-IDF值而不是很少出现的单词的LDA值。因此,我们构建了高度全面的日语词汇难度水平词典。我们使用构建的词典确认了测试数据集中所有单词的词汇水平都可以评分。
We are developing a Moodle plug-in, which is an AES (automated essay scoring) support system for the basic education of university students. Our system evaluates essays based on rubric, which has five evaluation viewpoints “Contents, Structure, Evidence, Style, and Skill”. Vocabulary level is one of the scoring items of Skill. It is calculated using Japanese Language Learners’ Dictionaries constructed by Sunakawa et al. Since this does not fully cover the words used in the student-level essays, we found that there is a problem with the accuracy of the vocabulary level scoring. In this paper, we propose to construct comprehensive Japanese vocabulary difficulty level dictionaries using Japanese Wikipedia as the corpus. We apply Latent Dirichlet Allocation (LDA) to the Wikipedia corpus and find the word appearance probability as one of the indexes of word difficulty. We use the TF-IDF value instead of the LDA value of the words, which rarely appears. As a result, we constructed highly comprehensive Japanese vocabulary difficulty level dictionaries. We confirmed that the vocabulary level can be scored for all words in the test dataset by using the constructed dictionaries.