Large language models to identify social determinants of health in electronic health records.

Large language models to identify social determinants of health in electronic health records.
复制标题

DOI:
10.1038/s41746-023-00970-0
复制
发表时间:
2024-01-11
影响因子:
15.2
通讯作者:
--
中科院分区:
医学1区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

健康的社会决定因素 (SDoH) 在患者治疗结果中发挥着至关重要的作用,但其记录在电子健康记录 (EHR) 的结构化数据中经常缺失或不完整。大型语言模型 (LLM) 可以从 EHR 中高通量提取 SDoH,以支持研究和临床护理。然而,类别不平衡和数据限制给这些记录稀疏但关键的信息带来了挑战。在这里,我们研究了使用法学硕士从 EHR 叙述文本中提取六个 SDoH 类别的最佳方法:就业、住房、交通、父母状况、关系和社会支持。性能最佳的模型是针对任何 SDoH 提及(宏观 F1 0.71)进行微调的 Flan-T5 XL,以及针对不利的 SDoH 提及(宏观 F1 0.70)进行微调的 Flan-T5 XXL。将 LLM 生成的合成数据添加到训练中因模型和架构而异,但提高了较小 Flan-T5 模型的性能(增量 F1 + 0.12 至 +0.23)。我们经过最佳微调的模型在零次和几次射击设置中的性能优于 ChatGPT 系列模型的零次和几次射击性能,但 GPT4 除外,它具有 10 次射击提示不良 SDoH 的功能。当将种族/民族和性别描述符添加到文本中时,微调模型比 ChatGPT 更不可能改变其预测,这表明算法偏差较小 (p<0.05)。我们的模型识别出 93.8% 的 SDoH 不良患者,而 ICD-10 代码识别出 2.0%。这些结果证明了法学硕士在改善 SDoH 的现实世界证据和帮助识别可以从资源支持中受益的患者方面的潜力。
Social determinants of health (SDoH) play a critical role in patient outcomes, yet their documentation is often missing or incomplete in the structured data of electronic health records (EHRs). Large language models (LLMs) could enable high-throughput extraction of SDoH from the EHR to support research and clinical care. However, class imbalance and data limitations present challenges for this sparsely documented yet critical information. Here, we investigated the optimal methods for using LLMs to extract six SDoH categories from narrative text in the EHR: employment, housing, transportation, parental status, relationship, and social support. The best-performing models were fine-tuned Flan-T5 XL for any SDoH mentions (macro-F1 0.71), and Flan-T5 XXL for adverse SDoH mentions (macro-F1 0.70). Adding LLM-generated synthetic data to training varied across models and architecture, but improved the performance of smaller Flan-T5 models (delta F1 + 0.12 to +0.23). Our best-fine-tuned models outperformed zero- and few-shot performance of ChatGPT-family models in the zero- and few-shot setting, except GPT4 with 10-shot prompting for adverse SDoH. Fine-tuned models were less likely than ChatGPT to change their prediction when race/ethnicity and gender descriptors were added to the text, suggesting less algorithmic bias (p < 0.05). Our models identified 93.8% of patients with adverse SDoH, while ICD-10 codes captured 2.0%. These results demonstrate the potential of LLMs in improving real-world evidence on SDoH and assisting in identifying patients who could benefit from resource support.
DOI: 10.1001/jama.2016.4226
发表时间: 2016-04-26
期刊: JAMA
影响因子: --
作者:
Chetty R;Stepner M;Abraham S;Lin S;Scuderi B;Turner N;Bergeron A;Cutler D
通讯作者: Cutler D
DOI: 10.1055/s-0040-1702214
发表时间: 2020-01-01
影响因子: 2.9
作者:
Feller, Daniel J.;Walk, Oliver J. Bear Don't;Elhadad, Noemie
通讯作者: Elhadad, Noemie
DOI: 10.3390/children1030390
发表时间: 2014-11-03
期刊: Children (Basel, Switzerland)
影响因子: --
作者:
Franke HA
通讯作者: Franke HA
DOI: 10.1038/s41551-021-00751-8
发表时间: 2021-06
影响因子: 28.1
作者:
通讯作者: --
DOI: 10.1093/jamia/ocx059
发表时间: 2018-01-01
影响因子: 6.4
作者:
Bejan, Cosmin A.;Angiolillo, John;Denny, Joshua C.
通讯作者: Denny, Joshua C.