Beyond Automated Essay Scoring Pioneering Research the Debate on Automated Essay Grading Recent Research Operational Writing-evaluation Systems Current Ets Writing Research Future Research and Applications

Beyond Automated Essay Scoring Pioneering Research the Debate on Automated Essay Grading Recent Research Operational Writing-evaluation Systems Current Ets Writing Research Future Research and Applications
复制标题

超越自动作文评分 开创性研究 关于自动作文评分的争论 最新研究 操作性写作评估系统 当前 Ets 写作研究 未来研究和应用

DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
L. Ferro
L. Ferro
中科院分区:
--
文献类型:
--
作者:
Marti A. Hearst;K. Kukich;L. Hirschman;Eric Breck;J. Burger;L. Ferro

文献摘要

参考文献

被引文献

相似文献

用自然语言交流的能力一直被认为是人类智能的一个定义性特征。此外,我们把用书面表达思想的能力看作是这种独特的人类语言能力的顶峰它蔑视公式化或算法规范。因此,毫不奇怪,试图设计计算机程序来评估写作的尝试经常遭到强烈的怀疑。尽管如此,自动化写作评估系统可能会提供我们所需要的平台,以阐明好的和坏的写作的许多特征,以及作为人类阅读和写作能力基础的许多语言、认知和其他技能。使用计算机来增加我们对文本特征的理解,以及创造和理解书面文本所涉及的认知技能,将有明显的好处。它将帮助我们开发更有效的教学材料,以提高阅读,写作和其他人类沟通能力。它还将帮助我们开发更有效的技术,如搜索引擎和问答系统,以提供普遍获得电子信息的机会。自动写作评估研究的简要历史及其未来方向的草图可能会为这一论点提供一些证据。Ellis Page为自动化写作评估搭建了舞台(参见图1中的时间轴)。[1]认识到在评估学生论文时对教师和大规模测试项目的巨大需求,佩奇开发了一个名为Project Essay Grader的自动论文评分系统。他从一组老师已经打分的学生作文开始。然后,他尝试了各种自动提取的文本特征,并应用多元线性回归来确定最佳的加权特征组合,以最好地预测教师的成绩。然后,他的系统可以使用相同的加权特征集为其他文章打分。PEG的分数与教师的分数呈0.78的多R相关性-几乎与两个或多个教师之间的0.85相关性一样强。在20世纪60年代,我们可以从文本中自动提取的特征仅限于表面特征。佩奇发现的一些最具预测性的特征包括平均单词长度、文章字数、逗号数量、介词数量和不常用词数量--后者与文章分数呈负相关。佩奇称这些特征代表了写作能力的某些内在品质。他不得不使用间接措施,因为计算困难的实施更直接的措施。尽管它在预测教师论文方面取得了令人印象深刻的成功...
The ability to communicate in natural language has long been considered a defining characteristic of human intelligence. Furthermore, we hold our ability to express ideas in writing as a pinnacle of this uniquely human language facility—it defies formulaic or algorithmic specification. So it comes as no surprise that attempts to devise computer programs that evaluate writing are often met with resounding skepticism. Nevertheless, automated writing-evaluation systems might provide precisely the platforms we need to elucidate many of the features that characterize good and bad writing, and many of the linguistic, cognitive, and other skills that underlie the human capacity for both reading and writing. Using computers to increase our understanding of the textual features and cognitive skills involved in creating and comprehending written text will have clear benefits. It will help us develop more effective instructional materials for improving reading, writing, and other human communication abilities. It will also help us develop more effective technologies , such as search engines and question-answering systems, for providing universal access to electronic information. A sketch of the brief history of automated writing-evaluation research and its future directions might lend some credence to this argument. Ellis Page set the stage for automated writing evaluation (see the timeline in Figure 1). 1 Recognizing the heavy demand placed on teachers and large-scale testing programs in evaluating student essays, Page developed an automated essay-grading system called Project Essay Grader. He started with a set of student essays that teachers had already graded. He then experimented with a variety of automatically extractable textual features and applied multiple linear regression to determine an optimal combination of weighted features that best predicted the teachers' grades. His system could then score other essays using the same set of weighted features. PEG's scores showed a multiple R correlation with teachers' scores of .78—almost as strong as the .85 correlation between two or more teachers. In the 1960s, the kinds of features we could automatically extract from text were limited to surface features. Some of the most predictive features Page found included average word length, essay length in words, number of commas, number of prepositions, and number of uncommon words—the latter being negatively correlated with essay scores. Page called these features proxies for some intrinsic qualities of writing competence. He had to use indirect measures because of the computational difficulty of implementing more direct measures. Despite its impressive success at predicting teachers' essay …
DOI: 10.1080/01638539809545028
发表时间: 1998-01-01
影响因子: 2.2
作者:
Landauer, TK;Foltz, PW;Laham, D
通讯作者: Laham, D