The reliability of workplace-based assessment in postgraduate medical education and training: a national evaluation in general practice in the United Kingdom

The reliability of workplace-based assessment in postgraduate medical education and training: a national evaluation in general practice in the United Kingdom
复制标题

DOI:
10.1007/s10459-008-9104-8
复制
发表时间:
2009-05-01
影响因子:
4
通讯作者:
Eva, Kevin W.
Eva, Kevin W.
中科院分区:
教育学2区
文献类型:
--
作者:
Murphy, Douglas J.;Bruce, David A.;Eva, Kevin W.

文献摘要

被引文献

相似文献

调查全科医学培训中六种潜在的基于工作场所的评估方法的可靠性和可行性:标准审核、临床和非临床同事的多源反馈、患者反馈(CARE 措施)、转介信、重大事件分析和咨询视频分析。在给定评估者和所需评估数量的情况下,使用每种工具对全科医生注册员(学员)的表现进行评估,以评估工具的可靠性和可行性。参与者对过程的体验由问卷确定。来自九个院长(代表英国所有四个国家)的 171 名全科医生注册员及其培训师参加了培训。使用普遍性理论评估每种工具区分医生的能力(可靠性)。然后进行决策研究,以确定使用每种仪器实现“高风险评估”可接受的高可靠性所需的观察数量。最后,使用描述性统计来总结参与者对使用这些工具的体验的评分。来自同事的多源反馈和患者的咨询反馈成为最有可能提供可靠且可行的工作场所绩效意见的两种方法。当进行两次评估时,通过 41 份 CARE Measure 患者调查问卷以及每位医生 6 名临床和/或 5 名非临床同事,可达到 0.8 的可靠性系数。对于其他四种测试方法,每个医生需要10名或更多评估员才能获得可靠的评估,这使得它们在高风险评估中使用的可行性极低。参与者的反馈并未对这些工具的可接受性、可行性或教育影响产生任何重大担忧。患者和同事对医生表现的看法相结合,再加上可靠的能力衡量标准,可以为监测全科医生培训的进展和完成情况提供合适的证据基础。
To investigate the reliability and feasibility of six potential workplace-based assessment methods in general practice training: criterion audit, multi-source feedback from clinical and non-clinical colleagues, patient feedback (the CARE Measure), referral letters, significant event analysis, and video analysis of consultations. Performance of GP registrars (trainees) was evaluated with each tool to assess the reliabilities of the tools and feasibility, given raters and number of assessments needed. Participant experience of process determined by questionnaire. 171 GP registrars and their trainers, drawn from nine deaneries (representing all four countries in the UK), participated. The ability of each tool to differentiate between doctors (reliability) was assessed using generalisability theory. Decision studies were then conducted to determine the number of observations required to achieve an acceptably high reliability for "high-stakes assessment" using each instrument. Finally, descriptive statistics were used to summarise participants' ratings of their experience using these tools. Multi-source feedback from colleagues and patient feedback on consultations emerged as the two methods most likely to offer a reliable and feasible opinion of workplace performance. Reliability co-efficients of 0.8 were attainable with 41 CARE Measure patient questionnaires and six clinical and/or five non-clinical colleagues per doctor when assessed on two occasions. For the other four methods tested, 10 or more assessors were required per doctor in order to achieve a reliable assessment, making the feasibility of their use in high-stakes assessment extremely low. Participant feedback did not raise any major concerns regarding the acceptability, feasibility, or educational impact of the tools. The combination of patient and colleague views of doctors' performance, coupled with reliable competence measures, may offer a suitable evidence-base on which to monitor progress and completion of doctors' training in general practice.