ChatGPT Performance on the American Urological Association Self-assessment Study Program and the Potential Influence of Artificial Intelligence in Urologic Training

ChatGPT Performance on the American Urological Association Self-assessment Study Program and the Potential Influence of Artificial Intelligence in Urologic Training
复制标题

DOI:
10.1016/j.urology.2023.05.010
复制
发表时间:
2023-08-10
期刊:
影响因子:
2.1
通讯作者:
Terlecki, Ryan
Terlecki, Ryan
中科院分区:
医学4区
文献类型:
--
作者:
Deebel, Nicholas A.;Terlecki, Ryan

文献摘要

被引文献

相似文献

目的 评估聊天生成预训练 Transformer (ChatGPT) 在美国泌尿外科协会自我评估研究计划 (AUA SASP) 中的表现,并按问题主干复杂性对表现进行分层。 方法 2021-2022 年 AUA SASP 计划中的问题由 ChatGPT 版本 3 (ChatGPT-3) 管理。使用标准化提示向模型提出问题。然后使用 ChatGPT 选择的答案选项来回答 AUA SASP 程序中的问题主干。然后,ChatGPT 会被提示为每个问题分配问题主干顺序(第一、第二、第三)。正确回答问题的百分比是针对每个订单级别确定的。 ChatGPT 提供的所有答复都经过定性评估,以了解适当的理由。 结果 ChatGPT 总共处理了 268 个问题。与 2022 年 AUA SASP 问题集相比,ChatGPT 在 2021 年的表现更好,正确回答了 42.3% 与 30.0% 的问题 (P < .05)。结论 ChatGPT 正确回答了许多高级问题,并为每个答案选择提供了合理的理由。虽然 ChatGPT 无法回答众多一阶问题,但未来的语言处理模型学习可能会导致其知识库的优化。这可能会导致利用像 ChatGPT 这样的人工智能作为泌尿科学员和教授的教育工具。泌尿学 177: 29-33, 2023.(c) 2023 Elsevier Inc. 保留所有权利。
OBJECTIVE To assess chat generative pre-trained transformer's (ChatGPT) performance on the American Urological Association Self-Assessment Study Program (AUA SASP) and stratify performance by question stem complexity.METHODS Questions from the 2021-2022 AUA SASP program were administered to ChatGPT version 3 (ChatGPT-3). Questions were administered to the model utilizing a standardized prompt. The answer choice selected by ChatGPT was then used to answer the question stem in the AUA SASP program. ChatGPT was then prompted to assign a question stem order (first, second, third) to each question. The percentage of correctly answered questions was determined for each order level. All responses provided by ChatGPT were qualitatively assessed for appropriate rationale.RESULTS A total of 268 questions were administered to ChatGPT. ChatGPT performed better on 2021 compared to the 2022 AUA SASP question set, answering 42.3% versus 30.0% of questions correctly (P .05).CONCLUSION ChatGPT answered many high-level questions correctly and provided a reasonable rationale for each answer choice. While ChatGPT was unable to answer numerous first-order questions, future language processing model learning may lead to the optimization of its fund of knowledge. This may lead to the utilization of artificial intelligence like ChatGPT as an educational tool for urology trainees and professors. UROLOGY 177: 29-33, 2023.(c) 2023 Elsevier Inc. All rights reserved.