Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of Its Successes and Shortcomings.

Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of Its Successes and Shortcomings.
复制标题

DOI:
10.1016/j.xops.2023.100324
复制
发表时间:
2023-12
影响因子:
--
通讯作者:
Duval, Renaud
Duval, Renaud
中科院分区:
其他
文献类型:
--
作者:
Antaki, Fares;Touma, Samir;Milad, Daniel;El -Khoury, Jonathan;Duval, Renaud

文献摘要

参考文献

被引文献

相似文献

基础模型是一种新型的人工智能算法,其中模型在未注释的数据上进行大规模预训练,并对无数的下游任务进行微调,例如生成文本。本研究评估了大型语言模型(LLM) ChatGPT在眼科问答领域的准确性。诊断试验或技术的评价。ChatGPT是一个公开可用的法学硕士。我们在两个流行的选择题库上测试了两个版本的ChatGPT(1月9日的“legacy”和ChatGPT Plus),这些选择题库通常用于准备高风险的眼科知识评估计划(OKAP)考试。我们从基础和临床科学课程(BCSC)自我评估计划和眼科问题在线题库中生成了两个260个问题的模拟考试。我们进行了逻辑回归来确定考试类别、认知水平和难度指数对答案准确性的影响。我们还使用Tukey检验进行了事后分析,以确定在测试的亚专业之间是否存在有意义的差异。我们通过比较ChatGPT的输出和题库提供的答案键来报告ChatGPT在每个考试部分的准确性,以正确率百分比表示。我们用似然比(LR)卡方给出了逻辑回归结果。我们认为各检查切片之间的差异在P值< 0.05时具有统计学意义。遗留模型在BCSC集上的准确率为55.8%,在OphthoQuestions集上的准确率为42.7%。使用ChatGPT Plus,准确率分别提高到59.4%±0.6%和49.2%±1.0%。在控制考试部分和认知水平的情况下,更容易的问题提高了准确性。遗留模型的Logistic回归分析表明,考试部分(LR, 27.57; P = 0.006)其次是问题难度(LR, 24.05; P < 0.001)最能预测ChatGPT的答案准确性。虽然传统模型在普通医学方面表现最好,在神经眼科和眼部病理学方面表现最差(P < 0.001),但ChatGPT Plus没有发现类似的事后发现,这表明在各个检查部分的结果更加一致。ChatGPT在模拟OKAP考试中的表现令人鼓舞。通过特定领域的预训练对法学硕士进行专业化培训可能是必要的,以提高他们在眼科亚专业的表现。在引用后可能会发现专有或商业披露。
Foundation models are a novel type of artificial intelligence algorithms, in which models are pretrained at scale on unannotated data and fine-tuned for a myriad of downstream tasks, such as generating text. This study assessed the accuracy of ChatGPT, a large language model (LLM), in the ophthalmology question-answering space. Evaluation of diagnostic test or technology. ChatGPT is a publicly available LLM. We tested 2 versions of ChatGPT (January 9 “legacy” and ChatGPT Plus) on 2 popular multiple choice question banks commonly used to prepare for the high-stakes Ophthalmic Knowledge Assessment Program (OKAP) examination. We generated two 260-question simulated exams from the Basic and Clinical Science Course (BCSC) Self-Assessment Program and the OphthoQuestions online question bank. We carried out logistic regression to determine the effect of the examination section, cognitive level, and difficulty index on answer accuracy. We also performed a post hoc analysis using Tukey’s test to decide if there were meaningful differences between the tested subspecialties. We reported the accuracy of ChatGPT for each examination section in percentage correct by comparing ChatGPT’s outputs with the answer key provided by the question banks. We presented logistic regression results with a likelihood ratio (LR) chi-square. We considered differences between examination sections statistically significant at a P value of < 0.05. The legacy model achieved 55.8% accuracy on the BCSC set and 42.7% on the OphthoQuestions set. With ChatGPT Plus, accuracy increased to 59.4% ± 0.6% and 49.2% ± 1.0%, respectively. Accuracy improved with easier questions when controlling for the examination section and cognitive level. Logistic regression analysis of the legacy model showed that the examination section (LR, 27.57; P = 0.006) followed by question difficulty (LR, 24.05; P < 0.001) were most predictive of ChatGPT’s answer accuracy. Although the legacy model performed best in general medicine and worst in neuro-ophthalmology (P < 0.001) and ocular pathology (P = 0.029), similar post hoc findings were not seen with ChatGPT Plus, suggesting more consistent results across examination sections. ChatGPT has encouraging performance on a simulated OKAP examination. Specializing LLMs through domain-specific pretraining may be necessary to improve their performance in ophthalmic subspecialties. Proprietary or commercial disclosure may be found after the references.
DOI: 10.1038/s41746-021-00464-x
发表时间: 2021-06-03
影响因子: 15.2
作者:
Korngiebel DM;Mooney SD
通讯作者: Mooney SD
DOI: 10.1136/bjophthalmol-2021-319030
发表时间: 2021-08-02
影响因子: 4.1
作者:
Antaki, Fares;Coussa, Razek Georges;Duval, Renaud
通讯作者: Duval, Renaud
DOI: 10.1186/s12909-019-1637-4
发表时间: 2019-06-07
影响因子: 3.6
作者:
Zafar, Sidra;Wang, Xueyang;Woreta, Fasika A.
通讯作者: Woreta, Fasika A.
DOI: 10.1016/j.ophtha.2012.06.010
发表时间: 2012-10-01
期刊: OPHTHALMOLOGY
影响因子: 13.7
作者:
Lee, Andrew G.;Oetting, Thomas A.;Zimmerman, M. Bridget
通讯作者: Zimmerman, M. Bridget
DOI: 10.2307/2529310
发表时间: 1977-01-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
LANDIS, JR;KOCH, GG
通讯作者: KOCH, GG