Beyond Automated Essay Scoring Pioneering Research the Debate on Automated Essay Grading Recent Research Operational Writing-evaluation Systems Current Ets Writing Research Future Research and Applications
Beyond Automated Essay Scoring Pioneering Research the Debate on Automated Essay Grading Recent Research Operational Writing-evaluation Systems Current Ets Writing Research Future Research and Applications
复制标题
超越自动作文评分 开创性研究 关于自动作文评分的争论 最新研究 操作性写作评估系统 当前 Ets 写作研究 未来研究和应用
DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
L. Ferro
中科院分区:
文献类型:
--
作者:
Marti A. Hearst;K. Kukich;L. Hirschman;Eric Breck;J. Burger;L. Ferro
The ability to communicate in natural language has long been considered a defining characteristic of human intelligence. Furthermore, we hold our ability to express ideas in writing as a pinnacle of this uniquely human language facility—it defies formulaic or algorithmic specification. So it comes as no surprise that attempts to devise computer programs that evaluate writing are often met with resounding skepticism. Nevertheless, automated writing-evaluation systems might provide precisely the platforms we need to elucidate many of the features that characterize good and bad writing, and many of the linguistic, cognitive, and other skills that underlie the human capacity for both reading and writing. Using computers to increase our understanding of the textual features and cognitive skills involved in creating and comprehending written text will have clear benefits. It will help us develop more effective instructional materials for improving reading, writing, and other human communication abilities. It will also help us develop more effective technologies , such as search engines and question-answering systems, for providing universal access to electronic information. A sketch of the brief history of automated writing-evaluation research and its future directions might lend some credence to this argument. Ellis Page set the stage for automated writing evaluation (see the timeline in Figure 1). 1 Recognizing the heavy demand placed on teachers and large-scale testing programs in evaluating student essays, Page developed an automated essay-grading system called Project Essay Grader. He started with a set of student essays that teachers had already graded. He then experimented with a variety of automatically extractable textual features and applied multiple linear regression to determine an optimal combination of weighted features that best predicted the teachers' grades. His system could then score other essays using the same set of weighted features. PEG's scores showed a multiple R correlation with teachers' scores of .78—almost as strong as the .85 correlation between two or more teachers. In the 1960s, the kinds of features we could automatically extract from text were limited to surface features. Some of the most predictive features Page found included average word length, essay length in words, number of commas, number of prepositions, and number of uncommon words—the latter being negatively correlated with essay scores. Page called these features proxies for some intrinsic qualities of writing competence. He had to use indirect measures because of the computational difficulty of implementing more direct measures. Despite its impressive success at predicting teachers' essay …
影响因子:
2.2
作者:
Landauer, TK;Foltz, PW;Laham, D
通讯作者:
Laham, D