Automated Scoring Using A Hybrid Feature Identification Technique

Automated Scoring Using A Hybrid Feature Identification Technique
复制标题

使用混合特征识别技术的自动评分

DOI:
--
复制
发表时间:
1998
期刊:
Annual Meeting of the Association for Computational Linguistics
影响因子:
--
通讯作者:
M. D. Harris
M. D. Harris
中科院分区:
--
文献类型:
--
作者:
J. Burstein;K. Kukich;Susanne Wolff;Chi Lu;M. Chodorow;Lisa C. Braden;M. D. Harris

文献摘要

被引文献

相似文献

这项研究利用自然语言固有的统计冗余来自动预测论文的分数。我们使用混合特征识别方法,包括句法结构分析、修辞结构分析和主题分析,对研究生管理入学考试(GMAT)和书面英语考试(TWE)考生的论文回答进行评分。对于每个论文问题,在训练集(人类得分论文回答的样本)上运行逐步线性回归分析,以提取每个测试问题的加权预测特征集。交叉验证集的分数预测是从预测特征集计算出来的。在15个测试题目中,电子作文评分器(e-rater)的预测分数与人类评分之间的准确或接近一致性在87%到94%之间。
This study exploits statistical redundancy inherent in natural language to automatically predict scores for essays. We use a hybrid feature identification method, including syntactic structure analysis, rhetorical structure analysis, and topical analysis, to score essay responses from test-takers of the Graduate Management Admissions Test (GMAT) and the Test of Written English (TWE). For each essay question, a stepwise linear regression analysis is run on a training set (sample of human scored essay responses) to extract a weighted set of predictive features for each test question. Score prediction for cross-validation sets is calculated from the set of predictive features. Exact or adjacent agreement between the Electronic Essay Rater (e-rater) score predictions and human rater scores ranged from 87% to 94% across the 15 test questions.