Applying item response theory (IRT) modeling to questionnaire development, evaluation, and refinement

Applying item response theory (IRT) modeling to questionnaire development, evaluation, and refinement
复制标题

DOI:
10.1007/s11136-007-9198-0
复制
发表时间:
2007-01-01
影响因子:
3.5
通讯作者:
Reeve, Bryce B.
Reeve, Bryce B.
中科院分区:
医学2区
文献类型:
--
作者:
Edelen, Maria Orlando;Reeve, Bryce B.

文献摘要

被引文献

相似文献

背景健康结果研究者越来越多地将项目反应理论(IRT)方法应用于问卷开发、评估和改进工作。目的提供IRT的简要概述,回顾与IRT应用相关的一些关键问题,并通过一个例子展示IRT的基本特征。504名青少年受访者在国家青少年健康纵向研究公共使用数据集谁完成了19项抑郁情绪量表。将样品分为开发和验证样品。在开发样本中使用分级反应模型校准量表项目,并将结果用于构建10项简表。结果19个条目的区分度不同,但在不同的条目中,被试的区分度不同,在不同的条目中,被试的区分度不同,在不同的条目中,被试的区分度不同(斜率参数范围:0.86 -2.66),项目位置参数反映了相当大的抑郁范围(-0.72-3.39)。然而,项目集是最歧视在较高水平的抑郁症。在验证样本中,短表和长表生成的IRT分数相关性为0.96,这些分数的平均差异为-0.01。此外,近90%的样本被归类为相同的风险或不处于抑郁症的风险,使用观察到的分数切割点从短期和长期forms.Conclusions当使用得当,IRT可以是一个强大的工具,问卷开发,评估和完善,导致精确,有效,相对简短的工具,最大限度地减少响应负担。
Background Health outcomes researchers are increasingly applying Item Response Theory (IRT) methods to questionnaire development, evaluation, and refinement efforts.Objective To provide a brief overview of IRT, to review some of the critical issues associated with IRT applications, and to demonstrate the basic features of IRT with an example.Methods Example data come from 6,504 adolescent respondents in the National Longitudinal Study of Adolescent Health public use data set who completed to the 19 item Feelings Scale for depression. The sample was split into a development and validation sample. Scale items were calibrated in the development sample with the Graded Response Model and the results were used to construct a 10-item short form. The short form was evaluated in the validation sample by examining the correspondence between IRT scores from the short form and the original, and by comparing the proportion of respondents identified as depressed according to the original and short form observed cut scores.Results The 19-items varied in their discrimination (slope parameter range: .86-2.66), and item location parameters reflected a considerable range of depression (-.72-3.39). However, the item set is most discriminating at higher levels of depression. In the validation sample IRT scores generated from the short and long forms were correlated at .96 and the average difference in these scores was -.01. In addition, nearly 90% of the sample was classified identically as at risk or not at risk for depression using observed score cut points from the short and long forms.Conclusions When used appropriately, IRT can be a powerful tool for questionnaire development, evaluation, and refinement, resulting in precise, valid, and relatively brief instruments that minimize response burden.