COMPUTER GRADING OF STUDENT PROSE, USING MODERN CONCEPTS AND SOFTWARE

COMPUTER GRADING OF STUDENT PROSE, USING MODERN CONCEPTS AND SOFTWARE
复制标题

DOI:
10.1080/00220973.1994.9943835
复制
发表时间:
1994-12-01
影响因子:
2.2
通讯作者:
PAGE, EB
PAGE, EB
中科院分区:
教育学4区
文献类型:
--
作者:
PAGE, EB

文献摘要

被引文献

相似文献

在早期的项目论文等级(PEG)的工作中,我们使用计算机来评估高中学生的散文。在主要的实验中,PEG成功地模仿了单个人类的评级,尽管20世纪60年代末的硬件和软件都很粗糙。今天,计算机在家庭和学校中很普遍,先进的软件包允许更强大的分析。在本研究中,我们分析了最近的联邦样本495和599论文和模拟组的人类法官,达到多个Rs高达0.87,接近目标法官组的明显可靠性。我们还从每个三分之二的形成样本中生成权重。其很好地预测了其它三分之一样本(简单RS高于0.84)。另一项交叉验证预测了不同年份、学生和评委小组的结果,r = 0.83。因此,计算机超过了两名法官,这是通常的人类小组。结果似乎令人鼓舞的进一步研究和早期应用到大型项目的论文评估和报告。
In earlier work of Project Essay Grade (PEG) we used computers to evaluate prose of high school students. In major experiments, PEG successfully imitated single human ratings, despite the crude hardware and software of the late 1960s. Today, computers are common in home and school, and advanced software packages permit much more powerful analysis. In the present research we analyzed recent federal samples of 495 and 599 essays and simulated groups of human judges, reaching multiple Rs as high as .87, close to the apparent reliability of the targeted judge groups. We also generated weights from formative samples of two thirds each. which predicted well the other one-third samples (with simple rs higher than .84). Another cross-validation predicted across different years, students, and judge panels, with an r of .83. Thus, the computer surpassed two judges, which is the usual human panel. Results appear encouraging for further research and indeed for early application to large programs of essay evaluation and reporting.