Algorithmic Identification of Treatment-Emergent Adverse Events From Clinical Notes Using Large Language Models: A Pilot Study in Inflammatory Bowel Disease.

Algorithmic Identification of Treatment-Emergent Adverse Events From Clinical Notes Using Large Language Models: A Pilot Study in Inflammatory Bowel Disease.
复制标题

使用大型语言模型从临床记录中算法识别治疗中出现的不良事件:炎症性肠病的初步研究。

DOI:
10.1002/cpt.3226
复制
发表时间:
2024
影响因子:
6.7
通讯作者:
B
B
中科院分区:
医学2区
文献类型:
--
作者:
Silverman,AnnaL;Sushil,Madhumita;Bhasuran,Balu;Ludwig,Dana;Buchanan,James;Racz,Rebecca;Parakala,Mahalakshmi;El-Kamary,Samer;Ahima,Ohenewaa;Belov,Artur;Choi,Lauren;Billings,Monisha;Li,Yan;Habal,Nadia;Liu,Qi;Tiwari,Jawahar;B

文献摘要

相似文献

门诊临床记录是关于药物安全性的丰富信息来源。然而,由于文本挖掘的方法学限制,这些注释中的数据目前未充分用于药物警戒。大型语言模型(LLM),如来自变压器的双向编码器表示(BERT),在一系列自然语言处理任务中取得了进展,但尚未在不良事件(AE)检测方面进行评估。我们采用了一项新的临床LLM(加州大学-旧金山弗朗西斯科(UCSF)-BERT),以确定使用非类固醇免疫抑制剂治疗炎症性肠病(IBD)后发生的严重AE(SAE)。我们将此模型与以前应用于AE检测的其他语言模型进行了比较。我们对928例IBD患者的928份门诊IBD记录进行了注释,记录了非类固醇免疫抑制剂治疗后发生的所有SAE相关住院治疗。这些记录共包含703起SAE,其中最常见的是预期疗效失败。在八个候选模型中,UCSF‐BERT在从该语料库中识别药物-SAE对方面实现了最高的数值性能(准确率88- 92%,宏F1 61-68%),比以前发表的模型准确率高5-10%。UCSF‐BERT在识别药物使用后出现的住院事件方面具有显著上级优势(P< 0.01)。LLM(如UCSF‐BERT)在从临床记录中检测SAE这一具有挑战性的任务上,与之前的方法相比,在数值上具有上级准确性。未来的工作需要适应这种方法,以提高模型性能和使用多中心数据和较新的架构,如生成预训练Transformer(GPT)的评估。我们的研究结果支持使用大型语言模型来加强药物警戒的潜在价值。
Outpatient clinical notes are a rich source of information regarding drug safety. However, data in these notes are currently underutilized for pharmacovigilance due to methodological limitations in text mining. Large language models (LLMs) like Bidirectional Encoder Representations from Transformers (BERT) have shown progress in a range of natural language processing tasks but have not yet been evaluated on adverse event (AE) detection. We adapted a new clinical LLM, University of California – San Francisco (UCSF)‐BERT, to identify serious AEs (SAEs) occurring after treatment with a non‐steroid immunosuppressant for inflammatory bowel disease (IBD). We compared this model to other language models that have previously been applied to AE detection. We annotated 928 outpatient IBD notes corresponding to 928 individual patients with IBD for all SAE‐associated hospitalizations occurring after treatment with a non‐steroid immunosuppressant. These notes contained 703 SAEs in total, the most common of which was failure of intended efficacy. Out of eight candidate models, UCSF‐BERT achieved the highest numerical performance on identifying drug‐SAE pairs from this corpus (accuracy 88–92%, macro F1 61–68%), with 5–10% greater accuracy than previously published models. UCSF‐BERT was significantly superior at identifying hospitalization events emergent to medication use (P< 0.01). LLMs like UCSF‐BERT achieve numerically superior accuracy on the challenging task of SAE detection from clinical notes compared with prior methods. Future work is needed to adapt this methodology to improve model performance and evaluation using multicenter data and newer architectures like Generative pre‐trained transformer (GPT). Our findings support the potential value of using large language models to enhance pharmacovigilance.