Extracting Family History Information From Electronic Health Records: Natural Language Processing Analysis.

Extracting Family History Information From Electronic Health Records: Natural Language Processing Analysis.
复制标题

DOI:
10.2196/24020
复制
发表时间:
2021-04-30
影响因子:
3.2
通讯作者:
Nguyen A
Nguyen A
中科院分区:
医学3区
文献类型:
--
作者:
Rybinski M;Dai X;Singh S;Karimi S;Nguyen A

文献摘要

参考文献

被引文献

相似文献

如果患者的家族史(FH)是已知的,则许多遗传性疾病和家族性疾病的预后、诊断和治疗显著改善。这些信息通常写在临床笔记的自由文本中。本研究的目的是开发自动化的方法,使访问FH数据通过自然语言处理。我们通过使用transformers从笔记中提取疾病提及来进行信息提取。我们还尝试了基于规则的方法,用于从文本和共指消解技术中提取家庭成员(FM)信息。我们评估了不同的迁移学习策略,以提高疾病的注释。我们提供了一个彻底的错误分析的影响因素,这样的信息提取系统。我们的实验表明,当在国家自然语言处理临床挑战赛(N2 C2)的公共共享任务数据集上进行测试时,领域自适应预训练和中间任务预训练的组合在从笔记中提取疾病和FM方面取得了81.63%的F1分数,与基线相比有统计学显著改善(P<.001)。相比之下,在2019年N2 C2/Open Health自然语言处理共享任务中,所有17支参赛队伍的F1得分中位数为76.59%。我们的方法利用最先进的命名实体识别模型进行疾病提及检测,再加上FM提及检测的混合方法,实现了接近参与2019年N2 C2 FH提取挑战的前3个系统的有效性,只有顶级系统在精度方面令人信服地优于我们的方法。
The prognosis, diagnosis, and treatment of many genetic disorders and familial diseases significantly improve if the family history (FH) of a patient is known. Such information is often written in the free text of clinical notes. The aim of this study is to develop automated methods that enable access to FH data through natural language processing. We performed information extraction by using transformers to extract disease mentions from notes. We also experimented with rule-based methods for extracting family member (FM) information from text and coreference resolution techniques. We evaluated different transfer learning strategies to improve the annotation of diseases. We provided a thorough error analysis of the contributing factors that affect such information extraction systems. Our experiments showed that the combination of domain-adaptive pretraining and intermediate-task pretraining achieved an F1 score of 81.63% for the extraction of diseases and FMs from notes when it was tested on a public shared task data set from the National Natural Language Processing Clinical Challenges (N2C2), providing a statistically significant improvement over the baseline (P<.001). In comparison, in the 2019 N2C2/Open Health Natural Language Processing Shared Task, the median F1 score of all 17 participating teams was 76.59%. Our approach, which leverages a state-of-the-art named entity recognition model for disease mention detection coupled with a hybrid method for FM mention detection, achieved an effectiveness that was close to that of the top 3 systems participating in the 2019 N2C2 FH extraction challenge, with only the top system convincingly outperforming our approach in terms of precision.
DOI: 10.2196/21750
发表时间: 2020-12-01
影响因子: 3.2
作者:
Dai HJ;Lee YQ;Nekkantti C;Jonnagaddala J
通讯作者: Jonnagaddala J
DOI: 10.1093/bioinformatics/btz682
发表时间: 2020-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lee J;Yoon W;Kim S;Kim D;Kim S;So CH;Kang J
通讯作者: Kang J
DOI: 10.1038/srep26094
发表时间: 2016-05-17
期刊: Scientific reports
影响因子: 4.6
作者:
Miotto R;Li L;Kidd BA;Dudley JT
通讯作者: Dudley JT
DOI: 10.1016/j.jbi.2015.07.010
发表时间: 2015-10
影响因子: 4.5
作者:
Leaman R;Khare R;Lu Z
通讯作者: Lu Z
DOI: 10.1016/j.jbi.2013.12.006
发表时间: 2014-02
影响因子: 4.5
作者:
Dogan, Rezarta Islamaj;Leaman, Robert;Lu, Zhiyong
通讯作者: Lu, Zhiyong