Reliability: on the reproducibility of assessment data

Reliability: on the reproducibility of assessment data
复制标题

DOI:
10.1111/j.1365-2929.2004.01932.x
复制
发表时间:
2004-09-01
期刊:
影响因子:
6
通讯作者:
Downing, SM
Downing, SM
中科院分区:
教育学1区
文献类型:
--
作者:
Downing, SM

文献摘要

被引文献

相似文献

所有评估数据,像其他科学实验数据一样,必须是可重复的,以便有意义地解释。目的探讨信度在医学教育中最常用的评估方法中的应用。对典型的可靠性估计方法进行了直观和非数学的讨论。信度是指评估结果的一致性。最重要的一致性的确切类型取决于评估的类型、目的和数据的相应使用。认知成就的笔试着眼于内部测试的一致性,使用从重测设计中衍生出来的估计方法。基于评分的评估数据,如对病房临床表现的评分,需要评分者之间的一致性或一致性。客观结构化的临床检查、模拟病人检查和其他性能类型的评估通常需要通用性理论分析,以解释复杂设计中测量误差的各种来源,并估计通用性对某个领域或技能的一致性。结论信度是评估效度证据的主要来源。低信度表明,在重新测试时,分数可能会有很大的变化。不一致的评估分数很难或不可能有意义地解释,从而降低了有效性证据。可靠性系数允许量化和估计评估中测量的随机误差,从而可以改进总体评估。
CONTEXT All assessment data, like other scientific experimental data, must be reproducible in order to be meaningfully interpreted.PURPOSE The purpose of this paper is to discuss applications of reliability to the most common assessment methods in medical education. Typical methods of estimating reliability are discussed intuitively and non-mathematically.SUMMARY Reliability refers to the consistency of assessment outcomes. The exact type of consistency of greatest interest depends on the type of assessment, its purpose and the consequential use of the data. Written tests of cognitive achievement look to internal test consistency, using estimation methods derived from the test-retest design. Rater-based assessment data, such as ratings of clinical performance on the wards, require interrater consistency or agreement. Objective structured clinical examinations, simulated patient examinations and other performance-type assessments generally require generalisability theory analysis to account for various sources of measurement error in complex designs and to estimate the consistency of the generalisations to a universe or domain of skills.CONCLUSIONS Reliability is a major source of validity evidence for assessments. Low reliability indicates that large variations in scores can be expected upon retesting. Inconsistent assessment scores are difficult or impossible to interpret meaningfully and thus reduce validity evidence. Reliability coefficients allow the quantification and estimation of the random errors of measurement in assessments, such that overall assessment can be improved.