Automated Essay Scoring at Scale: A Case Study in Switzerland and Germany

Automated Essay Scoring at Scale: A Case Study in Switzerland and Germany
复制标题

DOI:
10.1002/ets2.12249
复制
发表时间:
2019-03
影响因子:
--
通讯作者:
A. Rupp;J. Casabianca;Maleika Krüger;Stefan D. Keller;O. Köller
A. Rupp;J. Casabianca;Maleika Krüger;Stefan D. Keller;O. Köller
中科院分区:
--
文献类型:
--
作者:
A. Rupp;J. Casabianca;Maleika Krüger;Stefan D. Keller;O. Köller

文献摘要

被引文献

相似文献

在这份研究报告中,我们描述了一项大规模的论文写作能力研究的设计和实证结果,该研究涉及德国和瑞士约2,500名高中生,基于2个任务和2个相关提示,每个任务都来自标准化写作评估,其评分涉及人工和自动化组件。对于人工评分方面,我们描述了培训和监控人工评分员以及在定制平台内收集他们的评分的方法。对于自动评分方面,我们描述了训练,评估和选择适当的自动评分模型的方法,以及由此产生的任务分数与二级措施的分数的相关模式。分析表明,人类评分非常可靠,并且可以使用最先进的功能和机器学习方法构建有效的特定于机器人的自动评分模型,从而产生符合一般预期的二级指标的相关模式。最后,我们讨论了未来大规模开展此类工作的方法学意义。
In this research report, we describe the design and empirical findings for a large-scale study of essay writing ability with approximately 2,500 high school students in Germany and Switzerland on the basis of 2 tasks with 2 associated prompts, each from a standardized writing assessment whose scoring involved both human and automated components. For the human scoring aspect, we describe the methodology for training and monitoring human raters as well as for collecting their ratings within a customized platform. For the automated scoring aspect, we describe the methodology for training, evaluating, and selecting appropriate automated scoring models as well as correlational patterns of resulting task scores with scores from secondary measures. Analyses show that the human ratings were highly reliable and that effective prompt-specific automated scoring models could be built with state-of-the-art features and machine learning methods, which resulted in correlational patterns with secondary measures that were in line with general expectations. In closing, we discuss the methodological implications for conducting this kind of work at scale in the future.