ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?

ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?
复制标题

DOI:
10.1093/jamia/ocad104
复制
发表时间:
2023-06-19
影响因子:
6.4
通讯作者:
Lehmann, Christoph U.
Lehmann, Christoph U.
中科院分区:
管理学2区
文献类型:
--
作者:
Kumah-Crystal, Yaa;Mankowitz, Scott;Lehmann, Christoph U.

文献摘要

被引文献

相似文献

我们旨在评估ChatGPT在临床信息学委员会考试中的表现,并讨论大型语言模型(LLM)对委员会认证和维护的影响。我们使用Mankowitz的临床信息学委员会评论书中的260个多项选择题测试了ChatGPT,省略了6个与图像相关的问题。ChatGPT正确回答了254个合格问题中的190个(74%)。虽然临床信息学核心内容领域的性能各不相同,但差异无统计学意义。ChatGPT的表现引发了人们对医疗认证中潜在滥用和知识评估考试有效性的担忧。由于ChatGPT能够准确回答多项选择题,允许考生使用人工智能(AI)系统进行考试将损害家庭评估的可信度和有效性,并破坏公众的信任。AI和LLM的出现可能会颠覆现有的董事会认证和维护流程,并需要新的方法来评估医学教育的熟练程度。
We aimed to assess ChatGPT's performance on the Clinical Informatics Board Examination and to discuss the implications of large language models (LLMs) for board certification and maintenance. We tested ChatGPT using 260 multiple-choice questions from Mankowitz's Clinical Informatics Board Review book, omitting 6 image-dependent questions. ChatGPT answered 190 (74%) of 254 eligible questions correctly. While performance varied across the Clinical Informatics Core Content Areas, differences were not statistically significant. ChatGPT's performance raises concerns about the potential misuse in medical certification and the validity of knowledge assessment exams. Since ChatGPT is able to answer multiple-choice questions accurately, permitting candidates to use artificial intelligence (AI) systems for exams will compromise the credibility and validity of at-home assessments and undermine public trust. The advent of AI and LLMs threatens to upend existing processes of board certification and maintenance and necessitates new approaches to the evaluation of proficiency in medical education.