Zero-shot interpretable phenotyping of postpartum hemorrhage using large language models.

Zero-shot interpretable phenotyping of postpartum hemorrhage using large language models.
复制标题

DOI:
10.1038/s41746-023-00957-x
复制
发表时间:
2023-11-30
影响因子:
15.2
通讯作者:
--
中科院分区:
医学1区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

许多医学领域将受益于更深入,更准确的表型,但有有限的方法,表型使用临床笔记没有大量的注释数据。大型语言模型(LLM)已经证明了巨大的潜力,可以通过指定特定于任务的指令来适应新的任务,而无需额外的训练。在这里,我们报告了公开提供的法学硕士Flan-T5在使用电子健康记录中的出院记录对产后出血(PPH)患者进行表型分析方面的表现(n = 271,081)。该语言模型在提取与PPH相关的24个粒度概念方面表现出色。准确地识别这些颗粒概念可以开发可解释的复杂表型和亚型。Flan-T5模型在PPH的表型分型方面达到了高度的保真度(阳性预测值为0.95),与目前使用索赔代码的标准相比,识别出了47%以上的患有这种并发症的患者。这种LLM管道可以可靠地用于PPH亚型分型,并且在与子宫收缩乏力、异常胎盘形成和产科创伤相关的三种最常见的PPH亚型上优于基于索赔的方法。这种亚型分型方法的优点是其可解释性,因为有助于亚型确定的每个概念都可以评估。此外,由于新的指导方针,定义可能会随着时间的推移而变化,因此使用粒度概念来创建复杂的表型可以快速有效地更新算法。使用这种语言建模方法可以快速进行表型分析,而无需在多个临床用例中手动注释任何训练数据。
Many areas of medicine would benefit from deeper, more accurate phenotyping, but there are limited approaches for phenotyping using clinical notes without substantial annotated data. Large language models (LLMs) have demonstrated immense potential to adapt to novel tasks with no additional training by specifying task-specific instructions. Here we report the performance of a publicly available LLM, Flan-T5, in phenotyping patients with postpartum hemorrhage (PPH) using discharge notes from electronic health records (n = 271,081). The language model achieves strong performance in extracting 24 granular concepts associated with PPH. Identifying these granular concepts accurately allows the development of interpretable, complex phenotypes and subtypes. The Flan-T5 model achieves high fidelity in phenotyping PPH (positive predictive value of 0.95), identifying 47% more patients with this complication compared to the current standard of using claims codes. This LLM pipeline can be used reliably for subtyping PPH and outperforms a claims-based approach on the three most common PPH subtypes associated with uterine atony, abnormal placentation, and obstetric trauma. The advantage of this approach to subtyping is its interpretability, as each concept contributing to the subtype determination can be evaluated. Moreover, as definitions may change over time due to new guidelines, using granular concepts to create complex phenotypes enables prompt and efficient updating of the algorithm. Using this language modelling approach enables rapid phenotyping without the need for any manually annotated training data across multiple clinical use cases.
DOI: 10.1038/nbt.2749
发表时间: 2013-12
影响因子: 46.9
作者:
通讯作者: --
DOI: 10.1371/journal.pone.0192360
发表时间: 2018
期刊: PloS one
影响因子: 3.7
作者:
Gehrmann S;Dernoncourt F;Li Y;Carlson ET;Wu JT;Welt J;Foote J Jr;Moseley ET;Grant DW;Tyler PD;Celi LA
通讯作者: Celi LA
DOI: 10.1002/pds.4967
发表时间: 2020-03-02
影响因子: 2.6
作者:
He, Mengdong;Huybrechts, Krista F.;Bateman, Brian T.
通讯作者: Bateman, Brian T.
DOI: 10.1097/aog.0000000000004972
发表时间: 2023-01-01
影响因子: 7.2
作者:
Corbetta-Rastelli, Chiara M.;Friedman, Alexander M.;Wen, Timothy
通讯作者: Wen, Timothy
DOI: 10.1007/s10995-007-0256-6
发表时间: 2008-07-01
影响因子: 2.3
作者:
Kuklina, Elena V.;Whiteman, Maura K.;Marchbanks, Polly A.
通讯作者: Marchbanks, Polly A.