Dimensionality and Generalizability of Domain-Independent Performance Assessments

Dimensionality and Generalizability of Domain-Independent Performance Assessments
复制标题

与领域无关的绩效评估的维度和普遍性

DOI:
--
复制
发表时间:
1996
期刊:
影响因子:
--
通讯作者:
D. Niemi
D. Niemi
中科院分区:
--
文献类型:
--
作者:
E. Baker;J. Abedi;R. Linn;D. Niemi

文献摘要

被引文献

相似文献

可比较绩效评估设计的经验指导非常缺乏。一项研究进行了评估的程度,领域规格控制的主题和评分员的变异性,侧重于任务的概括性,评分员的可靠性,和评分规则的维度。两个历史班的学生被管理三个按需,多步骤的性能任务,一个星期分开。对于每个主题,所有学生都完成了先验知识测试,阅读主要源材料,并写了一篇解释性文章。使用基于理论的评分规则,四名训练有素的评分员对所有文章进行评分。报告了大鼠间和大鼠内的信度和g-研究结果。结果表明,相对效率的评估方法。维度分析支持两个因素:三个主题的深层理解和表层理解。历史课程的先验知识分数和GPA与评分规则的深度理解元素相关。对设计和测试的影响…
Abstract Empirical guidance for the design of comparable performance assessments is sorely lacking. A study was conducted to assess the degree to which domain specifications control topic and rater variability, focusing on task generalizability, rater reliability, and scoring rubric dimensionality. Two classes of history students were administered three on-demand, multistep performance tasks a week apart. For each topic, all students completed a Prior Knowledge Test, read primary source materials, and wrote an essay of explanation. Using a theory-based scoring rubric, four trained raters scored all essays. Inter- and intrarater reliabilities and g-study results are reported. Results show relative efficiency for the assessment approach. The dimensionality analysis supported two factors: Deep Understanding and Surface Understanding across the three topics. Prior Knowledge scores and GPA in history courses correlated with the Deep Understanding elements of the scoring rubric. Implications for design and test...