Assessment of writing ability in secondary education: comparison of analytic and holistic scoring systems for use in large-scale assessments

Assessment of writing ability in secondary education: comparison of analytic and holistic scoring systems for use in large-scale assessments
复制标题

中等教育写作能力评估:大规模评估中使用的分析评分系统和整体评分系统的比较

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
K. Böhme
K. Böhme
中科院分区:
--
文献类型:
--
作者:
S. Schipolowski;K. Böhme

文献摘要

被引文献

相似文献

虽然写作是中学语文教学的重要科目,但在大规模的评估中,写作往往被忽视。我们报道了一项在国家教育成绩监测的背景下对1365名8年级德国高中生进行的研究结果。学生对七个不同的说服性和信息性写作任务的反应使用了两种不同的评分系统:(I)分析性评分,14个二分标准,包括内容、文本结构和语言使用的具体方面;(Ii)整体评分,基于类似NAEP整体评分指南的综合评分表,并伴随着内容、风格(即语言使用和组织)和语言正确性的半整体评分。我们考察了两种评分程序的评分结果,包括评分人之间和评分人内的可靠性、维度和评分结果的趋同性。调查结果在不同的写作任务和语篇体裁中的概括性也受到了关注。结果表明,整体和半整体量表比大多数分析标准具有更好的信度。对于这两种评分系统,内容和结构方面密切相关,而语言正确性是一个明显不同的维度。两种评分系统都测量了相同的潜在结构。
Although writing is an important subject of language teaching in secondary education, it is often neglected in large-scale assessments. We report results of a study with 1,365 German high school students in Grade 8 that was conducted in the context of national monitoring of educational achievement. Student responses on seven different persuasive and informative writing tasks were evaluated with two different scoring systems: (i) analytic scoring with 14 dichotomous criteria capturing specific aspects of content, text structure, and language usage, and (ii) holistic scoring based on a comprehensive rating scale similar to the NAEP Holistic Scoring Guide accompanied by semi-holistic scales for content, style (i.e., language usage and organization), and language correct- ness. We inspected the results of both scoring procedures in terms of inter-rater and intra-rater reliability, dimensionality, and convergence of the scoring results. Attention is also given to the generalizability of the findings across different writing tasks and text genres. The results showed better reliability for the holistic and semi-holistic scales than for most of the analytic criteria. For both scoring systems, content and structural aspects were closely associated whereas language correctness was a clearly distinct dimension. Both scoring systems measured the same latent construct.