Improving large language models for clinical named entity recognition via prompt engineering.

Improving large language models for clinical named entity recognition via prompt engineering.
复制标题

DOI:
10.1093/jamia/ocad259
复制
发表时间:
2023-03
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Yan Hu;Iqra Ameer;X. Zuo;Xueqing Peng;Yujia Zhou;Zehan Li;Yiming Li;Jianfu Li;Xiaoqian Jiang;Hua Xu
Yan Hu;Iqra Ameer;X. Zuo;Xueqing Peng;Yujia Zhou;Zehan Li;Yiming Li;Jianfu Li;Xiaoqian Jiang;Hua Xu
中科院分区:
其他
文献类型:
--
作者:
Yan Hu;Iqra Ameer;X. Zuo;Xueqing Peng;Yujia Zhou;Zehan Li;Yiming Li;Jianfu Li;Xiaoqian Jiang;Hua Xu

文献摘要

被引文献

相似文献

这项研究强调了大型语言模型,特别是GPT-3.5和GPT-4在处理复杂的临床数据和用最少的训练数据提取有意义的信息方面的潜力。通过开发和改进基于提示的策略,我们可以显著提高模型的性能,使其成为临床NER任务的可行工具,并可能减少对大量注释数据集的依赖。目的量化GPT-3.5和GPT-4在临床命名实体识别(NER)任务中的能力,并提出针对任务的提示以提高其性能。材料与方法我们在两个临床NER任务上对这些模型进行了评估:(1)在2010年i2b2概念提取共享任务之后,从MTSamples语料库的临床笔记中提取医疗问题、治疗和测试,以及(2)从疫苗不良事件报告系统(VAERS)的安全报告中识别与神经系统疾病相关的不良事件。为了提高GPT模型的性能,我们开发了一个临床任务特定提示框架,该框架包括(1)带有任务描述和格式规范的基线提示,(2)基于注释指南的提示,(3)基于错误分析的提示,以及(4)用于少量学习的注释样本。我们评估了每个提示的有效性,并将这些模型与BioClinicalBERT进行了比较。结果在基线提示下,GPT-3.5和GPT-4获得了放松的F1分数,MTSamples为0.634,0.804,VAERS为0.301,0.593。其他提示组件持续提高了模型性能。当使用所有4个组件时,GPT-3.5和GPT-4获得了松弛的F1 SOC值,MTSamples为0.794,0.861,VAERS为0.676,0.736,证明了我们的Prompt框架的有效性。尽管这些结果落后于BioClinicalBERT(MTSamples数据集的F1为0.901,VAERS为0.802),但考虑到需要的训练样本很少,它是非常有希望的。讨论这项研究的结果表明,在利用LLMS进行临床NER任务方面,有一个很有希望的方向。然而,虽然GPT模型的性能通过特定于任务的提示得到了改进,但仍需要进一步的开发和改进。像GPT-4这样的LLM在实现与BioClinicalBERT等最先进模型接近的性能方面具有潜力,但它们仍然需要仔细及时的工程设计和对特定任务知识的理解。这项研究还强调了评估模式的重要性,该模式准确地反映了低成本管理在临床环境中的能力和表现。结论虽然直接将GPT模型应用于临床NER任务并不能达到最佳性能,但我们的特定任务提示框架结合了医学知识和训练样本,显著提高了GPT模型在潜在临床应用中的可行性。
IMPORTANCE The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets. OBJECTIVES This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance. MATERIALS AND METHODS We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, and tests from clinical notes in the MTSamples corpus, following the 2010 i2b2 concept extraction shared task, and (2) to identify nervous system disorder-related adverse events from safety reports in the vaccine adverse event reporting system (VAERS). To improve the GPT models' performance, we developed a clinical task-specific prompt framework that includes (1) baseline prompts with task description and format specification, (2) annotation guideline-based prompts, (3) error analysis-based instructions, and (4) annotated samples for few-shot learning. We assessed each prompt's effectiveness and compared the models to BioClinicalBERT. RESULTS Using baseline prompts, GPT-3.5 and GPT-4 achieved relaxed F1 scores of 0.634, 0.804 for MTSamples and 0.301, 0.593 for VAERS. Additional prompt components consistently improved model performance. When all 4 components were used, GPT-3.5 and GPT-4 achieved relaxed F1 socres of 0.794, 0.861 for MTSamples and 0.676, 0.736 for VAERS, demonstrating the effectiveness of our prompt framework. Although these results trail BioClinicalBERT (F1 of 0.901 for the MTSamples dataset and 0.802 for the VAERS), it is very promising considering few training samples are needed. DISCUSSION The study's findings suggest a promising direction in leveraging LLMs for clinical NER tasks. However, while the performance of GPT models improved with task-specific prompts, there's a need for further development and refinement. LLMs like GPT-4 show potential in achieving close performance to state-of-the-art models like BioClinicalBERT, but they still require careful prompt engineering and understanding of task-specific knowledge. The study also underscores the importance of evaluation schemas that accurately reflect the capabilities and performance of LLMs in clinical settings. CONCLUSION While direct application of GPT models to clinical NER tasks falls short of optimal performance, our task-specific prompt framework, incorporating medical knowledge and training samples, significantly enhances GPT models' feasibility for potential clinical applications.