The future of outcomes measurement: item banking, tailored short-forms, and computerized adaptive assessment

The future of outcomes measurement: item banking, tailored short-forms, and computerized adaptive assessment
复制标题

DOI:
10.1007/s11136-007-9204-6
复制
发表时间:
2007-01-01
影响因子:
3.5
通讯作者:
Choi, Seung
Choi, Seung
中科院分区:
医学2区
文献类型:
--
作者:
Cella, David;Gershon, Richard;Choi, Seung

文献摘要

被引文献

相似文献

使用题库和计算机化自适应测验(CAT)首先明确定义重要结果,并将这些定义引用到收集到大量经过充分研究的题库或“题库”中的具体问题。项目可以从库中选择以形成定制的短度表,或者可以由为精确度和临床相关性而编程的计算机确定的顺序和长度进行管理。虽然远不完美,但这样的题库可以形成对人类症状和功能问题的共同定义和理解,如疲劳、疼痛、抑郁、行动能力、社会功能、感觉功能和许多其他我们只能通过直接询问人们才能衡量的健康概念。美国国立卫生研究院(NIH)通过名为PROMIS(www.nihpromis.org)的NIH路线图倡议(www.nihpromis.org)与测量专家达成的合作协议证明了这一点,这是朝着这个方向迈出的一大步。我们对题库和CAT的方法是实用的;无论是科学还是理论,我们都专注于应用。从实践的角度来看,我们经常必须决定是重写和重新测试一个项目,增加更多的项目来填补空白(通常在衡量标准的上限),在一些修改后重新测试银行,或者将银行拆分成更单一、但临床相关性更低或更完整的单元。这些决定并不容易,但它们很少是不可饶恕的。我们鼓励人们建立实用的工具,能够从普通银行产生多个短表衡量标准和CAT管理,并加深我们对这些具有不同临床人口和年龄的银行的理解,以便随着时间的推移,从这些许多活动中产生的分数不仅开始具有共同的衡量标准和范围,而且开始具有用户之间的共享意义和理解。在本文中,我们提供了题库和CAT的概述,讨论了我们的题库及其副产品的方法,描述了测试选项,讨论了CAT用于疲劳的例子,并讨论了诸如PROIS这样的实体的长期可持续性的模型。成功的一些障碍包括方法本身的限制,方法之间的争议和分歧,以及最终用户不愿放弃熟悉的方法。
The use of item banks and computerized adaptive testing (CAT) begins with clear definitions of important outcomes, and references those definitions to specific questions gathered into large and well-studied pools, or "banks" of items. Items can be selected from the bank to form customized short scales, or can be adminis-tered in a sequence and length determined by a computer programmed for precision and clinical relevance. Although far from perfect, such item banks can form a common definition and understanding of human symptoms and functional problems such as fatigue, pain, depression, mobility, social function, sensory function, and many other health concepts that we can only measure by asking people directly. The support of the National Institutes of Health (NIH), as witnessed by its cooperative agreement with measurement experts through the NIH Roadmap Initiative known as PROMIS (www.nihpromis.org), is a big step in that direction. Our approach to item banking and CAT is practical; as focused on application as it is on science or theory. From a practical perspective, we frequently must decide whether to re-write and retest an item, add more items to fill gaps (often at the ceiling of the measure), retest a bank after some modifications, or split up a bank into units that are more unidimensional, yet less clinically relevant or complete. These decisions are not easy, and yet they are rarely unforgiving. We encourage people to build practical tools that are capable of producing multiple short form measures and CAT administrations from common banks, and to further our understanding of these banks with various clinical populations and ages, so that with time the scores that emerge from these many activities begin to have not only a common metric and range, but a shared meaning and understanding across users. In this paper, we provide an overview of item banking and CAT, discuss our approach to item banking and its byproducts, describe testing options, discuss an example of CAT for fatigue, and discuss models for long term sustainability of an entity such as PROMIS. Some barriers to success include limitations in the methods themselves, controversies and disagreements across approaches, and end-user reluctance to move away from the familiar.