A marker-based neural network system for extracting social determinants of health

A marker-based neural network system for extracting social determinants of health
复制标题

用于提取健康社会决定因素的基于标记的神经网络系统

DOI:
10.1093/jamia/ocad041
复制
发表时间:
2023
影响因子:
6.4
通讯作者:
Rios, Anthony
Rios, Anthony
中科院分区:
管理学2区
文献类型:
--
作者:
Zhao, Xingmeng;Rios, Anthony

文献摘要

参考文献

相似文献

健康的社会决定因素(SDoH)对患者的医疗保健质量和差异的影响是众所周知的。许多SDoH项目在电子健康记录中没有以结构化形式编码。这些项目通常在自由文本临床笔记中捕获,但自动提取它们的方法有限。我们探索一个多阶段的管道,涉及命名实体识别(NER),关系分类(RC),文本分类方法,自动提取SDoH信息从临床notes.Materials和MethodsThe研究使用N2 C2共享任务数据,这是从2个来源的临床笔记:MIMIC-III和华盛顿港景大学医学中心。它包含4480个社会历史部分,并为12个SDoH提供完整的注释。为了处理实体重叠的问题,我们开发了一种新的基于标记的NER模型。我们用它在一个多阶段的管道中提取SDoH信息从临床notes.ResultsOur标记为基础的系统优于国家的最先进的跨度为基础的模型在处理重叠的实体的基础上的整体微F1评分性能。与共享任务方法相比,它还实现了最先进的性能。我们的方法子任务A,B和C,分别实现了0.9101,0.8053和0.9025的F1。ConclusionsThe主要发现这项研究是,多阶段管道有效地提取SDoH信息从临床笔记。这种方法可以改善在临床环境中对SDoH的理解和跟踪。然而,错误传播可能是一个问题,需要进一步的研究,以提高复杂的语义和低频实体的实体提取。我们已经在https://github.com/Zephyr1022/SDOH-N2C2-UTSA上提供了源代码。
ObjectiveThe impact of social determinants of health (SDoH) on patients’ healthcare quality and the disparity is well known. Many SDoH items are not coded in structured forms in electronic health records. These items are often captured in free-text clinical notes, but there are limited methods for automatically extracting them. We explore a multi-stage pipeline involving named entity recognition (NER), relation classification (RC), and text classification methods to automatically extract SDoH information from clinical notes.Materials and MethodsThe study uses the N2C2 Shared Task data, which were collected from 2 sources of clinical notes: MIMIC-III and University of Washington Harborview Medical Centers. It contains 4480 social history sections with full annotation for 12 SDoHs. In order to handle the issue of overlapping entities, we developed a novel marker-based NER model. We used it in a multi-stage pipeline to extract SDoH information from clinical notes.ResultsOur marker-based system outperformed the state-of-the-art span-based models at handling overlapping entities based on the overall Micro-F1 score performance. It also achieved state-of-the-art performance compared with the shared task methods. Our approach achieved an F1 of 0.9101, 0.8053, and 0.9025 for Subtasks A, B, and C, respectively.ConclusionsThe major finding of this study is that the multi-stage pipeline effectively extracts SDoH information from clinical notes. This approach can improve the understanding and tracking of SDoHs in clinical settings. However, error propagation may be an issue and further research is needed to improve the extraction of entities with complex semantic meanings and low-frequency entities. We have made the source code available at https://github.com/Zephyr1022/SDOH-N2C2-UTSA.
DOI: 10.1377/hlthaff.2009.0730
发表时间: 2010-03-01
期刊: HEALTH AFFAIRS
影响因子: 9.7
作者:
Singh, Gopal K.;Siahpush, Mohammad;Kogan, Michael D.
通讯作者: Kogan, Michael D.
DOI: 10.1055/s-0040-1702214
发表时间: 2020-01-01
影响因子: 2.9
作者:
Feller, Daniel J.;Walk, Oliver J. Bear Don't;Elhadad, Noemie
通讯作者: Elhadad, Noemie
粮食不安全与糖尿病之间的交叉点:回顾
DOI: --
发表时间: 2014
影响因子: 4.9
作者:
E. Gucciardi;M. Vahabi;Nicole Norris;J. Del Monte;Cecile M. Farnum
通讯作者: Cecile M. Farnum
DOI: 10.1186/s13326-019-0198-0
发表时间: 2019-04-11
影响因子: 1.9
作者:
Conway, Mike;Keyhani, Salomeh;Chapman, Wendy W.
通讯作者: Chapman, Wendy W.