Automatic Difficulty Classification of Arabic Sentences

Automatic Difficulty Classification of Arabic Sentences
复制标题

阿拉伯语句子自动难度分类

DOI:
--
复制
发表时间:
2021
期刊:
Workshop on Arabic Natural Language Processing
影响因子:
--
通讯作者:
S. Sharoff
S. Sharoff
中科院分区:
--
文献类型:
--
作者:
Nouran Khallaf;S. Sharoff

文献摘要

参考文献

被引文献

相似文献

在本文中,我们提出了一种现代标准阿拉伯语 (MSA) 句子难度分类器,它使用 CEFR 熟练程度或简单或复杂的二元分类来预测语言学习者的句子难度。我们比较了不同类型的句子嵌入(fastText、mBERT、XLM-R 和阿拉伯语-BERT)的使用,以及传统语言特征,如 POS 标签、依存树、可读性分数和语言学习者的频率列表。我们的最佳结果是使用经过微调的阿拉伯语 BERT 取得的。我们的 3 向 CEFR 分类的准确度对于阿拉伯语-Bert 和 XLM-R 分类分别为 0.80 和 0.75 的 F-1,对于回归的 Spearman 相关性为 0.71。我们的二元难度分类器对于句子对语义相似度分类器达到 F-1 0.94 和 F-1 0.98。
In this paper, we present a Modern Standard Arabic (MSA) Sentence difficulty classifier, which predicts the difficulty of sentences for language learners using either the CEFR proficiency levels or the binary classification as simple or complex. We compare the use of sentence embeddings of different kinds (fastText, mBERT , XLM-R and Arabic-BERT), as well as traditional language features such as POS tags, dependency trees, readability scores and frequency lists for language learners. Our best results have been achieved using fined-tuned Arabic-BERT. The accuracy of our 3-way CEFR classification is F-1 of 0.80 and 0.75 for Arabic-Bert and XLM-R classification respectively and 0.71 Spearman correlation for regression. Our binary difficulty classifier reaches F-1 0.94 and F-1 0.98 for sentence-pair semantic similarity classifier.
阿拉伯语学习者语料库 v1:阿拉伯语研究的新资源
DOI: --
发表时间: 2013
期刊: --
影响因子: --
作者:
Alfaifi AYG
通讯作者: Alfaifi AYG