Acquisition of a Lexicon for Family History Information: Bidirectional Encoder Representations From Transformers-Assisted Sublanguage Analysis.

Acquisition of a Lexicon for Family History Information: Bidirectional Encoder Representations From Transformers-Assisted Sublanguage Analysis.
复制标题

获得家族史信息的词典获取:来自变形金刚辅助的串联分析的双向编码器表示。

DOI:
10.2196/48072
复制
发表时间:
2023-06-27
影响因子:
3.2
通讯作者:
Liu, Hongfang
Liu, Hongfang
中科院分区:
医学3区
文献类型:
--
作者:
Wang, Liwei;He, Huan;Wen, Andrew;Moon, Sungrim;Fu, Sunyang;Peterson, Kevin J.;Ai, Xuguang;Liu, Sijia;Kavuluru, Ramakanth;Liu, Hongfang

文献摘要

参考文献

相似文献

患者的家族史(FH)信息显着影响下游的临床护理。尽管如此重要,但还没有标准化的方法来捕获电子健康记录中的 FH 信息,并且很大一部分 FH 信息经常嵌入在临床记录中。这使得 FH 信息难以在下游数据分析或临床决策支持应用中使用。为了解决这个问题,可以使用能够提取和标准化跳频信息的自然语言处理系统。在这项研究中,我们的目的是构建一个用于信息提取和规范化的 FH 词汇资源。我们利用基于变压器的方法来构建 FH 词汇资源,利用由初级保健的一部分生成的临床笔记组成的语料库。该词典的可用性通过开发基于规则的 FH 系统得到了证明,该系统提取之前 FH 挑战中指定的 FH 实体和关系。我们还尝试了基于深度学习的跳频系统来提取跳频信息。之前的 FH 挑战数据集用于评估。生成的词典包含 33,603 个词典条目,标准化为统一医学语言系统的 6408 个概念唯一标识符和 15,126 个医学临床术语系统命名法代码,每个概念平均有 5.4 个变体。性能评估表明基于规则的跳频系统取得了合理的性能。基于规则的 FH 系统与最先进的基于深度学习的 FH 系统相结合,可以提高使用 BioCreative/N2C2 FH 挑战数据集评估的 FH 信息的召回率,F1 分数各不相同但具有可比性。由此产生的词典和基于规则的 FH 系统可通过 Open Health 自然语言处理 GitHub 免费获取。
A patient’s family history (FH) information significantly influences downstream clinical care. Despite this importance, there is no standardized method to capture FH information in electronic health records and a substantial portion of FH information is frequently embedded in clinical notes. This renders FH information difficult to use in downstream data analytics or clinical decision support applications. To address this issue, a natural language processing system capable of extracting and normalizing FH information can be used. In this study, we aimed to construct an FH lexical resource for information extraction and normalization. We exploited a transformer-based method to construct an FH lexical resource leveraging a corpus consisting of clinical notes generated as part of primary care. The usability of the lexicon was demonstrated through the development of a rule-based FH system that extracts FH entities and relations as specified in previous FH challenges. We also experimented with a deep learning–based FH system for FH information extraction. Previous FH challenge data sets were used for evaluation. The resulting lexicon contains 33,603 lexicon entries normalized to 6408 concept unique identifiers of the Unified Medical Language System and 15,126 codes of the Systematized Nomenclature of Medicine Clinical Terms, with an average number of 5.4 variants per concept. The performance evaluation demonstrated that the rule-based FH system achieved reasonable performance. The combination of the rule-based FH system with a state-of-the-art deep learning–based FH system can improve the recall of FH information evaluated using the BioCreative/N2C2 FH challenge data set, with the F1 score varied but comparable. The resulting lexicon and rule-based FH system are freely available through the Open Health Natural Language Processing GitHub.
DOI: 10.1016/j.jbi.2020.103541
发表时间: 2020-10
影响因子: 4.5
作者:
Peterson KJ;Jiang G;Liu H
通讯作者: Liu H
DOI: 10.1136/jamia.1999.0060205
发表时间: 1999-05-01
影响因子: 6.4
作者:
Johnson, SB
通讯作者: Johnson, SB
DOI: 10.1111/cts.13463
发表时间: 2023-03
期刊: Clinical and translational science
影响因子: --
作者:
通讯作者: --
DOI: 10.1200/cci.22.00006
发表时间: 2022-07
影响因子: 4.2
作者:
Wang, Liwei;Fu, Sunyang;Wen, Andrew;Ruan, Xiaoyang;He, Huan;Liu, Sijia;Moon, Sungrim;Mai, Michelle;Riaz, Irbaz B.;Wang, Nan;Yang, Ping;Xu, Hua;Warner, Jeremy L.;Liu, Hongfang
通讯作者: Liu, Hongfang
DOI: 10.1186/s13073-020-00819-1
发表时间: 2021-01-07
期刊: Genome medicine
影响因子: 12.3
作者:
Bylstra Y;Lim WK;Kam S;Tham KW;Wu RR;Teo JX;Davila S;Kuan JL;Chan SH;Bertin N;Yang CX;Rozen S;Teh BT;Yeo KK;Cook SA;Jamuar SS;Ginsburg GS;Orlando LA;Tan P
通讯作者: Tan P