Automated language essay scoring systems: a literature review.

Automated language essay scoring systems: a literature review.
复制标题

DOI:
10.7717/peerj-cs.208
复制
发表时间:
2019
期刊:
PeerJ. Computer science
影响因子:
--
通讯作者:
Nassef M
Nassef M
中科院分区:
其他
文献类型:
--
作者:
Hussein MA;Hassan H;Nassef M

文献摘要

参考文献

被引文献

相似文献

作文是衡量考生语言能力的重要因素,然而,对作文或短文的评分在可靠性和时间上都是一个非常具有挑战性的过程。对客观和快速评分的需求已经提出了对可以针对特定提示自动评分作文问题的计算机系统的需求。自动论文评分(AES)系统用于通过使用自然语言处理(NLP)和机器学习技术来克服写作任务评分的挑战。本文的目的是回顾文献AES系统用于评分的论述题。我们已经使用Google Scholar、EBSCO和ERIC对现有文献进行了回顾,以搜索用英语撰写的论文的术语“AES”、“Automated Essay Scoring”、“Automated Essay Grading”或“Automatic Essay”。已经确定了两类:手工制作的功能和自动功能的AES系统。前一类系统与设计特征的质量密切相关。另一方面,后一类系统是基于自动学习的特点和关系的文章和它的分数没有任何手工制作的功能。我们回顾了这两类系统的系统主要焦点,系统中使用的技术,对训练数据的需求,教学应用(反馈系统),以及电子分数和人类分数之间的相关性。本文包括三个主要部分。首先,我们提出了一个结构化的文献综述可用的手工特征AES系统。其次,我们提出了一个结构化的文献综述可用的自动特征AES系统。最后,我们得出了一系列的讨论和结论。已经发现AES模型利用广泛的手动调整的浅层和深层语言特征。AES系统在减少劳动密集型标记活动,确保评分标准的一致应用以及确保评分的客观性方面具有许多优势。尽管已经实施了许多技术来改进AES系统,但是已经确定了三个主要挑战。挑战是缺乏评分者作为一个人的感觉,系统可能会被欺骗,给一篇文章比它应得的更低或更高的分数,以及评估想法和命题的创造性和评估其实用性的能力有限。许多技术仅用于解决前两个挑战。
Writing composition is a significant factor for measuring test-takers’ ability in any language exam. However, the assessment (scoring) of these writing compositions or essays is a very challenging process in terms of reliability and time. The need for objective and quick scores has raised the need for a computer system that can automatically grade essay questions targeting specific prompts. Automated Essay Scoring (AES) systems are used to overcome the challenges of scoring writing tasks by using Natural Language Processing (NLP) and machine learning techniques. The purpose of this paper is to review the literature for the AES systems used for grading the essay questions. We have reviewed the existing literature using Google Scholar, EBSCO and ERIC to search for the terms “AES”, “Automated Essay Scoring”, “Automated Essay Grading”, or “Automatic Essay” for essays written in English language. Two categories have been identified: handcrafted features and automatically featured AES systems. The systems of the former category are closely bonded to the quality of the designed features. On the other hand, the systems of the latter category are based on the automatic learning of the features and relations between an essay and its score without any handcrafted features. We reviewed the systems of the two categories in terms of system primary focus, technique(s) used in the system, the need for training data, instructional application (feedback system), and the correlation between e-scores and human scores. The paper includes three main sections. First, we present a structured literature review of the available Handcrafted Features AES systems. Second, we present a structured literature review of the available Automatic Featuring AES systems. Finally, we draw a set of discussions and conclusions. AES models have been found to utilize a broad range of manually-tuned shallow and deep linguistic features. AES systems have many strengths in reducing labor-intensive marking activities, ensuring a consistent application of scoring criteria, and ensuring the objectivity of scoring. Although many techniques have been implemented to improve the AES systems, three primary challenges have been identified. The challenges are lacking of the sense of the rater as a person, the potential that the systems can be deceived into giving a lower or higher score to an essay than it deserves, and the limited ability to assess the creativity of the ideas and propositions and evaluate their practicality. Many techniques have only been used to address the first two challenges.
DOI: 10.1016/0165-7836(94)90020-5
发表时间: 1994-02-01
期刊: FISHERIES RESEARCH
影响因子: 2.4
作者:
CROZIER, WW;KENNEDY, GJA
通讯作者: KENNEDY, GJA
DOI: 10.1080/00220973.1994.9943835
发表时间: 1994-12-01
影响因子: 2.2
作者:
PAGE, EB
通讯作者: PAGE, EB