How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment.

How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment.
复制标题

DOI:
10.2196/45312
复制
发表时间:
2023-02-08
影响因子:
3.6
通讯作者:
Chartash, David
Chartash, David
中科院分区:
其他
文献类型:
--
作者:
Gilson, Aidan;Safranek, Conrad W;Huang, Thomas;Socrates, Vimig;Chi, Ling;Taylor, Richard Andrew;Chartash, David

文献摘要

被引文献

相似文献

ChatGPT(Chat Generative Pre-trained Transformer)是一个拥有1750亿个参数的自然语言处理模型,可以对用户输入生成对话式响应。本研究旨在评估ChatGPT在美国医师执照考试第1步和第2步考试范围内的问题上的表现,以及分析用户可解释性的回答。我们使用了两组多项选择题来评估ChatGPT的表现,每组都有与步骤1和步骤2相关的问题。第一组来自AMBOSS,这是一个医学生常用的题库,它还提供了关于问题难度和相对于用户群的考试表现的统计数据。第二组是国家医学考试委员会(NBME)免费的120个问题。ChatGPT的性能与其他两个大型语言模型GPT-3和InstructGPT进行了比较。每个ChatGPT响应的文本输出通过3个定性指标进行评估:所选答案的逻辑依据,问题内部信息的存在以及问题外部信息的存在。在4个数据集中,AMBOSS-Step1,AMBOSS-Step2,NBME-Free-Step1和NBME-Free-Step2,ChatGPT分别实现了44%(44/100),42%(42/100),64.4%(56/87)和57.8%(59/102)的准确率。ChatGPT在所有数据集中的平均表现优于InstructGPT 8.15%,GPT-3的表现与随机机会相似。该模型表明,在AMBOSS-步骤1数据集中,随着问题难度的增加(P=.01),性能显著下降。我们发现,ChatGPT的答案选择的逻辑理由存在于NBME数据集的100%输出中。在所有问题中,96.8%(183/189)的问题提供了该问题的内部信息。在NBME-Free-Step1(P<.001)和NBME-Free-Step2(P=.001)数据集上,相对于正确答案,错误答案的问题外部信息的存在分别低44.5%和27%。ChatGPT标志着自然语言处理模型在医疗问题回答任务上的重大改进。通过在NBME-Free-Step-1数据集上执行大于60%的阈值,我们表明该模型达到了相当于三年级医学生的及格分数。此外,我们强调ChatGPT在大多数答案中提供逻辑和信息上下文的能力。这些事实结合在一起,为ChatGPT作为交互式医学教育工具的潜在应用提供了令人信服的案例,以支持学习。
Chat Generative Pre-trained Transformer (ChatGPT) is a 175-billion-parameter natural language processing model that can generate conversation-style responses to user input. This study aimed to evaluate the performance of ChatGPT on questions within the scope of the United States Medical Licensing Examination Step 1 and Step 2 exams, as well as to analyze responses for user interpretability. We used 2 sets of multiple-choice questions to evaluate ChatGPT’s performance, each with questions pertaining to Step 1 and Step 2. The first set was derived from AMBOSS, a commonly used question bank for medical students, which also provides statistics on question difficulty and the performance on an exam relative to the user base. The second set was the National Board of Medical Examiners (NBME) free 120 questions. ChatGPT’s performance was compared to 2 other large language models, GPT-3 and InstructGPT. The text output of each ChatGPT response was evaluated across 3 qualitative metrics: logical justification of the answer selected, presence of information internal to the question, and presence of information external to the question. Of the 4 data sets, AMBOSS-Step1, AMBOSS-Step2, NBME-Free-Step1, and NBME-Free-Step2, ChatGPT achieved accuracies of 44% (44/100), 42% (42/100), 64.4% (56/87), and 57.8% (59/102), respectively. ChatGPT outperformed InstructGPT by 8.15% on average across all data sets, and GPT-3 performed similarly to random chance. The model demonstrated a significant decrease in performance as question difficulty increased (P=.01) within the AMBOSS-Step1 data set. We found that logical justification for ChatGPT’s answer selection was present in 100% of outputs of the NBME data sets. Internal information to the question was present in 96.8% (183/189) of all questions. The presence of information external to the question was 44.5% and 27% lower for incorrect answers relative to correct answers on the NBME-Free-Step1 (P<.001) and NBME-Free-Step2 (P=.001) data sets, respectively. ChatGPT marks a significant improvement in natural language processing models on the tasks of medical question answering. By performing at a greater than 60% threshold on the NBME-Free-Step-1 data set, we show that the model achieves the equivalent of a passing score for a third-year medical student. Additionally, we highlight ChatGPT’s capacity to provide logic and informational context across the majority of answers. These facts taken together make a compelling case for the potential applications of ChatGPT as an interactive medical education tool to support learning.