Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models.

Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models.
复制标题

DOI:
10.1371/journal.pdig.0000198
复制
发表时间:
2023-02
期刊:
PLOS digital health
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

我们评估了一个名为ChatGPT的大型语言模型在美国医学执照考试(USMLE)上的表现,该考试包括三个考试:步骤1,步骤2CK和步骤3。ChatGPT在所有三项考试中的表现都达到或接近及格门槛,没有任何专门的培训或强化。此外,ChatGPT在其解释中表现出高度的一致性和洞察力。这些结果表明,大型语言模型可能有助于医学教育,并可能有助于临床决策。人工智能(AI)系统在改善医疗保健和健康结果方面具有巨大的潜力。因此,确保临床AI的开发以信任和可解释性原则为指导至关重要。将人工智能的医学知识与人类临床专家的知识进行比较是评估这些质量的关键第一步。为了实现这一目标,我们评估了ChatGPT(一种基于语言的AI)在美国医学执照考试(USMLE)上的表现。USMLE是一套三个标准化的专家级知识测试,这是在美国获得医疗许可证所必需的。我们发现ChatGPT的准确率达到或接近60%的通过阈值。作为第一个实现这一基准的公司,这标志着人工智能成熟的一个重要里程碑。令人印象深刻的是,ChatGPT能够在没有人类培训师的专业投入的情况下实现这一结果。此外,ChatGPT表现出可理解的推理和有效的临床见解,增强了信任和可解释性的信心。我们的研究表明,像ChatGPT这样的大型语言模型可能会在医学教育环境中帮助人类学习者,作为未来整合到临床决策中的前奏。
We evaluated the performance of a large language model called ChatGPT on the United States Medical Licensing Exam (USMLE), which consists of three exams: Step 1, Step 2CK, and Step 3. ChatGPT performed at or near the passing threshold for all three exams without any specialized training or reinforcement. Additionally, ChatGPT demonstrated a high level of concordance and insight in its explanations. These results suggest that large language models may have the potential to assist with medical education, and potentially, clinical decision-making. Artificial intelligence (AI) systems hold great promise to improve medical care and health outcomes. As such, it is crucial to ensure that the development of clinical AI is guided by the principles of trust and explainability. Measuring AI medical knowledge in comparison to that of expert human clinicians is a critical first step in evaluating these qualities. To accomplish this, we evaluated the performance of ChatGPT, a language-based AI, on the United States Medical Licensing Exam (USMLE). The USMLE is a set of three standardized tests of expert-level knowledge, which are required for medical licensure in the United States. We found that ChatGPT performed at or near the passing threshold of 60% accuracy. Being the first to achieve this benchmark, this marks a notable milestone in AI maturation. Impressively, ChatGPT was able to achieve this result without specialized input from human trainers. Furthermore, ChatGPT displayed comprehensible reasoning and valid clinical insights, lending increased confidence to trust and explainability. Our study suggests that large language models such as ChatGPT may potentially assist human learners in a medical education setting, as a prelude to future integration into clinical decision-making.