Impact of different scoring algorithms applied to multiple-mark survey items on outcome assessment: an in-field study on health-related knowledge

Impact of different scoring algorithms applied to multiple-mark survey items on outcome assessment: an in-field study on health-related knowledge
复制标题

应用于多标记调查项目的不同评分算法对结果评估的影响:健康相关知识的现场研究

DOI:
10.15167/2421-4248/jpmh2015.56.4.464
复制
发表时间:
2015
影响因子:
--
通讯作者:
D. Amicizia
D. Amicizia
中科院分区:
--
文献类型:
--
作者:
A. Domnich;D. Panatto;L. Arata;I. Bevilacqua;L. Apprato;R. Gasparini;D. Amicizia

文献摘要

被引文献

相似文献

导言。与健康有关的知识通常通过多项选择测试进行评估。在不同类型的格式中,研究人员可以选择使用多标记项目,即有一个以上的正确答案。虽然多重评分项目长期以来一直在学术环境中使用-有时很少或不确定的结果-很少有人知道这种格式在现场健康教育和促进研究的实施。方法.以中学生为研究对象,进行营养知识问卷调查,并进行单次讲座干预。采用八种不同的评分算法对答案进行评分,并从经典测试理论的角度进行分析。同样的调查被重新管理的学生样本,以评估他们的知识的短期变化。结果共分析了286份问卷。部分评分算法显示出更好的心理测量学特性比二分规则。特别是Ripkey提出的算法和平衡规则在多评分项目中表现出更好的内部一致性和相对效率。一个惩罚算法,其中标记的干扰因素的比例减去标记的正确答案是唯一一个突出了显着差异的表现之间的本地人和移民,可能是由于其略好的歧视能力。该算法还与干预前/干预后评分变化的最大效应量相关。讨论在健康教育与促进研究中,选择合适的评分规则不仅要考虑单个算法的心理测量学特性,而且要考虑研究目的和结果,因为评分规则在偏倚性、可靠性、难度、猜测敏感性和区分度方面存在差异。
Summary Introduction. Health-related knowledge is often assessed through multiple-choice tests. Among the different types of formats, researchers may opt to use multiple-mark items, i.e. with more than one correct answer. Although multiple-mark items have long been used in the academic setting – sometimes with scant or inconclusive results – little is known about the implementation of this format in research on in-field health education and promotion. Methods. A study population of secondary school students completed a survey on nutrition-related knowledge, followed by a single- lecture intervention. Answers were scored by means of eight different scoring algorithms and analyzed from the perspective of classical test theory. The same survey was re-administered to a sample of the students in order to evaluate the short-term change in their knowledge. Results. In all, 286 questionnaires were analyzed. Partial scoring algorithms displayed better psychometric characteristics than the dichotomous rule. In particular, the algorithm proposed by Ripkey and the balanced rule showed greater internal consistency and relative efficiency in scoring multiple-mark items. A penalizing algorithm in which the proportion of marked distracters was subtracted from that of marked correct answers was the only one that highlighted a significant difference in performance between natives and immigrants, probably owing to its slightly better discriminatory ability. This algorithm was also associated with the largest effect size in the pre-/post-intervention score change. Discussion. The choice of an appropriate rule for scoring multiple- mark items in research on health education and promotion should consider not only the psychometric properties of single algorithms but also the study aims and outcomes, since scoring rules differ in terms of biasness, reliability, difficulty, sensitivity to guessing and discrimination.