Diagnostic Accuracy of Differential-Diagnosis Lists Generated by Generative Pretrained Transformer 3 Chatbot for Clinical Vignettes with Common Chief Complaints: A Pilot Study.

Diagnostic Accuracy of Differential-Diagnosis Lists Generated by Generative Pretrained Transformer 3 Chatbot for Clinical Vignettes with Common Chief Complaints: A Pilot Study.
复制标题

DOI:
10.3390/ijerph20043378
复制
发表时间:
2023-02-15
影响因子:
--
通讯作者:
Shimizu, Taro
Shimizu, Taro
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hirosawa, Takanobu;Harada, Yukinori;Yokose, Masashi;Sakamoto, Tetsu;Kawamura, Ren;Shimizu, Taro

文献摘要

被引文献

相似文献

人工智能(AI)聊天机器人,包括生成式预训练变压器3 (GPT-3)聊天机器人(ChatGPT-3)产生的鉴别诊断的诊断准确性尚不清楚。本研究评估了ChatGPT-3对具有常见主诉的临床小插曲生成的鉴别诊断列表的准确性。普通内科医师为10种常见主诉创建临床病例、正确诊断和5种鉴别诊断。ChatGPT-3在10个鉴别诊断表中的正诊率为28/30(93.3%)。在5个鉴别诊断列表中,医生的正确诊断率仍优于ChatGPT-3 (98.3% vs. 83.3%, p = 0.03)。在最高诊断中,医生的正确诊断率也优于ChatGPT-3 (53.3% vs. 93.3%, p < 0.001)。在ChatGPT-3生成的10个鉴别诊断列表中,医师鉴别诊断一致性率为62/88(70.5%)。综上所述,本研究证明了ChatGPT-3生成的鉴别诊断清单对具有常见主诉的临床病例具有较高的诊断准确性。这表明,像ChatGPT-3这样的人工智能聊天机器人可以为常见的主诉生成一份区分良好的诊断清单。然而,这些列表的顺序可以在未来得到改进。
The diagnostic accuracy of differential diagnoses generated by artificial intelligence (AI) chatbots, including the generative pretrained transformer 3 (GPT-3) chatbot (ChatGPT-3) is unknown. This study evaluated the accuracy of differential-diagnosis lists generated by ChatGPT-3 for clinical vignettes with common chief complaints. General internal medicine physicians created clinical cases, correct diagnoses, and five differential diagnoses for ten common chief complaints. The rate of correct diagnosis by ChatGPT-3 within the ten differential-diagnosis lists was 28/30 (93.3%). The rate of correct diagnosis by physicians was still superior to that by ChatGPT-3 within the five differential-diagnosis lists (98.3% vs. 83.3%, p = 0.03). The rate of correct diagnosis by physicians was also superior to that by ChatGPT-3 in the top diagnosis (53.3% vs. 93.3%, p < 0.001). The rate of consistent differential diagnoses among physicians within the ten differential-diagnosis lists generated by ChatGPT-3 was 62/88 (70.5%). In summary, this study demonstrates the high diagnostic accuracy of differential-diagnosis lists generated by ChatGPT-3 for clinical cases with common chief complaints. This suggests that AI chatbots such as ChatGPT-3 can generate a well-differentiated diagnosis list for common chief complaints. However, the order of these lists can be improved in the future.