Adapted large language models can outperform medical experts in clinical text summarization.

Adapted large language models can outperform medical experts in clinical text summarization.
复制标题

DOI:
10.1038/s41591-024-02855-5
复制
发表时间:
2023-09
期刊:
影响因子:
82.9
通讯作者:
Dave Van Veen;Cara Van Uden;Louis Blankemeier;Jean-Benoit Delbrouck;Asad Aali;Christian Blüthgen;A. Pareek;Malgorzata Polacin;William Collins;Neera Ahuja;C. Langlotz;Jason Hom;S. Gatidis;John M. Pauly;Akshay S. Chaudhari
Dave Van Veen;Cara Van Uden;Louis Blankemeier;Jean-Benoit Delbrouck;Asad Aali;Christian Blüthgen;A. Pareek;Malgorzata Polacin;William Collins;Neera Ahuja;C. Langlotz;Jason Hom;S. Gatidis;John M. Pauly;Akshay S. Chaudhari
中科院分区:
医学1区
文献类型:
--
作者:
Dave Van Veen;Cara Van Uden;Louis Blankemeier;Jean-Benoit Delbrouck;Asad Aali;Christian Blüthgen;A. Pareek;Malgorzata Polacin;William Collins;Neera Ahuja;C. Langlotz;Jason Hom;S. Gatidis;John M. Pauly;Akshay S. Chaudhari

文献摘要

被引文献

相似文献

分析大量的文本数据和总结电子健康记录中的关键信息对临床医生如何分配他们的时间造成了巨大的负担。尽管大型语言模型(LLM)在自然语言处理(NLP)任务中表现出了希望,但它们在各种临床摘要任务中的有效性尚未得到证实。在这里,我们应用自适应方法,八个LLM,跨越四个不同的临床总结任务:放射学报告,病人的问题,进度记录和医患对话。定量评估与句法,语义和概念NLP指标揭示模型和适应方法之间的权衡。一项由10名医生参与的临床阅片人研究评估了总结的完整性、正确性和简洁性;在大多数情况下,与医学专家的总结相比,我们最适合的LLM的总结被认为是等同的(45%)或上级的(36%)。随后的安全分析强调了法学硕士和医学专家所面临的挑战,因为我们将错误与潜在的医疗危害联系起来,并对捏造的信息进行分类。我们的研究提供了LLM在多个任务中的临床文本摘要方面优于医学专家的证据。这表明将LLM集成到临床工作流程中可以减轻文档负担,使临床医生能够更多地关注患者护理。
Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language processing (NLP) tasks, their effectiveness on a diverse range of clinical summarization tasks remains unproven. Here we applied adaptation methods to eight LLMs, spanning four distinct clinical summarization tasks: radiology reports, patient questions, progress notes and doctor–patient dialogue. Quantitative assessments with syntactic, semantic and conceptual NLP metrics reveal trade-offs between models and adaptation methods. A clinical reader study with 10 physicians evaluated summary completeness, correctness and conciseness; in most cases, summaries from our best-adapted LLMs were deemed either equivalent (45%) or superior (36%) compared with summaries from medical experts. The ensuing safety analysis highlights challenges faced by both LLMs and medical experts, as we connect errors to potential medical harm and categorize types of fabricated information. Our research provides evidence of LLMs outperforming medical experts in clinical text summarization across multiple tasks. This suggests that integrating LLMs into clinical workflows could alleviate documentation burden, allowing clinicians to focus more on patient care.