A Natural Language Processing Tool to Extract Quantitative Smoking Status from Clinical Narratives.

A Natural Language Processing Tool to Extract Quantitative Smoking Status from Clinical Narratives.
复制标题

DOI:
10.1109/ichi48887.2020.9374369
复制
发表时间:
2020-11
期刊:
Proceedings. IEEE International Conference on Healthcare Informatics
影响因子:
--
通讯作者:
Wu Y
Wu Y
中科院分区:
其他
文献类型:
--
作者:
Yang X;Yang H;Lyu T;Yang S;Guo Y;Bian J;Xu H;Wu Y

文献摘要

被引文献

相似文献

这项研究提出了一种自然语言处理(NLP)工具来提取定量吸烟信息(例如,包年,戒烟年,吸烟年,每天包),并将其标准化为包年单位。我们注释了200个临床笔记的语料库,这些笔记来自进行低剂量CT成像程序进行肺癌筛查的患者,并使用两层规则引擎结构开发了一个NLP系统。我们将200个音符分成训练集和测试集,并仅使用训练集开发NLP系统。在测试集上的实验结果表明,我们的NLP系统实现了最好的F1分数为0.963和0.946,分别为宽松和严格的评价。
This study presents a natural language processing (NLP) tool to extract quantitative smoking information (e.g., Pack-Year, Quit Year, Smoking Year, and Pack per Day) from clinical notes and standardized them into Pack-Year unit. We annotated a corpus of 200 clinical notes from patients who had low-dose CT imaging procedures for lung cancer screening and developed an NLP system using a two-layer rule-engine structure. We divided the 200 notes into a training set and a test set and developed the NLP system only using the training set. The experimental results on the test set showed that our NLP system achieved the best F1 scores of 0.963 and 0.946 for lenient and strict evaluation, respectively.