Large language models encode clinical knowledge.

Large language models encode clinical knowledge.
复制标题

DOI:
10.1038/s41586-023-06291-2
复制
发表时间:
2023-08
期刊:
影响因子:
64.8
通讯作者:
Natarajan, Vivek
Natarajan, Vivek
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Singhal, Karan;Azizi, Shekoofeh;Tu, Tao;Mahdavi, S. Sara;Wei, Jason;Chung, Hyung Won;Scales, Nathan;Tanwani, Ajay;Cole-Lewis, Heather;Pfohl, Stephen;Payne, Perry;Seneviratne, Martin;Gamble, Paul;Kelly, Chris;Babiker, Abubakr;Schaerli, Nathanael;Chowdhery, Aakanksha;Mansfield, Philip;Demner-Fushman, Dina;Arcas, Blaise Aguera y;Webster, Dale;Corrado, Greg S.;Matias, Yossi;Chou, Katherine;Gottweis, Juraj;Tomasev, Nenad;Liu, Yun;Rajkomar, Alvin;Barral, Joelle;Semturs, Christopher;Karthikesalingam, Alan;Natarajan, Vivek

文献摘要

参考文献

被引文献

相似文献

大型语言模型(LLM)已经展示了令人印象深刻的功能,但临床应用的门槛很高。评估模型的临床知识的尝试通常依赖于基于有限基准的自动评估。在这里,为了解决这些局限性,我们提出了MultiMedQA,一个基准,结合了六个现有的医疗问题回答数据集,涵盖专业医学,研究和消费者查询,以及一个新的在线搜索医疗问题数据集,HealthSearchQA。我们提出了一个人的评价框架模型的答案沿着多个轴,包括真实性,理解,推理,可能的伤害和偏见。此外,我们评估路径语言模型(PaLM,一个540亿参数LLM)和它的解释调整的变体,Flan-PaLM的MultiMedQA。通过结合使用提示策略,Flan-PaLM在每个MultiMedQA多项选择数据集(MedQA,MedMCQA,PubMedQA和测量大规模多任务语言理解(MMLU)临床主题)上都达到了最先进的准确率,其中包括MedQA(美国医疗许可考试式问题)的67.6%准确率,超过了现有技术的17%以上。然而,人类评估揭示了关键的差距。为了解决这个问题,我们引入了指令提示调整,一个参数有效的方法对齐LLM使用几个范例到新的域。由此产生的模型,Med-PaLM,表现良好,但仍然不如临床医生。我们发现,理解,知识回忆和推理提高模型规模和指令提示调整,这表明在医学LLM的潜在效用。我们的人类评估揭示了当今模型的局限性,加强了评估框架和方法开发在为临床应用创建安全,有用的LLM方面的重要性。Med-PaLM是一种最先进的医学大型语言模型,在几个医学问答任务中进行了介绍和评估,展示了这些模型在这一领域的前景。
Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We propose a human evaluation framework for model answers along multiple axes including factuality, comprehension, reasoning, possible harm and bias. In addition, we evaluate Pathways Language Model (PaLM, a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA, MedMCQA, PubMedQA and Measuring Massive Multitask Language Understanding (MMLU) clinical topics), including 67.6% accuracy on MedQA (US Medical Licensing Exam-style questions), surpassing the prior state of the art by more than 17%. However, human evaluation reveals key gaps. To resolve this, we introduce instruction prompt tuning, a parameter-efficient approach for aligning LLMs to new domains using a few exemplars. The resulting model, Med-PaLM, performs encouragingly, but remains inferior to clinicians. We show that comprehension, knowledge recall and reasoning improve with model scale and instruction prompt tuning, suggesting the potential utility of LLMs in medicine. Our human evaluations reveal limitations of today’s models, reinforcing the importance of both evaluation frameworks and method development in creating safe, helpful LLMs for clinical applications. Med-PaLM, a state-of-the-art large language model for medicine, is introduced and evaluated across several medical question answering tasks, demonstrating the promise of these models in this domain.
DOI: 10.1038/s41746-020-00376-2
发表时间: 2021-01-08
影响因子: 15.2
作者:
Esteva A;Chou K;Yeung S;Naik N;Madani A;Mottaghi A;Liu Y;Topol E;Dean J;Socher R
通讯作者: Socher R
DOI: 10.1038/s41581-021-00501-8
发表时间: 2022-03
期刊: Nature reviews. Nephrology
影响因子: --
作者:
Eneanya ND;Boulware LE;Tsai J;Bruce MA;Ford CL;Harris C;Morales LS;Ryan MJ;Reese PP;Thorpe RJ Jr;Morse M;Walker V;Arogundade FA;Lopes AA;Norris KC
通讯作者: Norris KC
DOI: 10.1162/tacl_a_00317
发表时间: 2020-01-01
影响因子: 10.9
作者:
Clark, Jonathan H.;Choi, Eunsol;Palomaki, Jennimaria
通讯作者: Palomaki, Jennimaria
DOI: 10.1016/j.aiopen.2022.11.003
发表时间: 2022-01-01
期刊: AI OPEN
影响因子: --
作者:
Han, Xu;Zhao, Weilin;Sun, Maosong
通讯作者: Sun, Maosong
DOI: 10.1146/annurev-biodatasci-092820-114757
发表时间: 2021-07
期刊: Annual review of biomedical data science
影响因子: --
作者:
Chen IY;Pierson E;Rose S;Joshi S;Ferryman K;Ghassemi M
通讯作者: Ghassemi M