Mining heart disease risk factors in clinical text with named entity recognition and distributional semantic models.

Mining heart disease risk factors in clinical text with named entity recognition and distributional semantic models.
复制标题

使用命名实体识别和分布式语义模型挖掘临床文本中的心脏病风险因素。

DOI:
10.1016/j.jbi.2015.08.009
复制
发表时间:
2015-12
影响因子:
4.5
通讯作者:
Urbain J
Urbain J
中科院分区:
医学3区
文献类型:
--
作者:
Urbain J

文献摘要

被引文献

相似文献

我们设计并分析了一个多阶段自然语言处理系统的性能,该系统采用命名实体识别、贝叶斯统计和规则逻辑来识别和表征糖尿病患者的心脏病危险因素事件。该系统最初是为2014 i2b2临床数据自然语言挑战而开发的。该系统的优势包括在识别与心脏病风险因素事件相关的命名实体方面具有很高的准确性。该系统的主要弱点是在描述某些事件的属性时不准确。例如,确定事件相对于记录日期的相对时间,事件是否可归因于患者的病史或患者的家族史,以及区分当前和以前的吸烟状况。我们认为,这些不准确性在很大程度上是由于缺乏将上下文集成到我们的事件检测模型中的有效方法。为了解决这些不准确的问题,我们探索添加一个分布式语义模型来表征心脏病危险因素事件的上下文证据。使用这个语义模型,我们将最初的2014 i2b2 Challenges in Natural Language of Clinical data F1得分从0.838提高到0.890,并且在不使用任何可能导致结果偏差的词汇的情况下,将精度提高了10.3%。用分布语义模型提取糖尿病的力向图和cad概念。
We present the design, and analyze the performance of a multi-stage natural language processing system employing named entity recognition, Bayesian statistics, and rule logic to identify and characterize heart disease risk factor events in diabetic patients over time. The system was originally developed for the 2014 i2b2 Challenges in Natural Language in Clinical Data. The system's strengths included a high level of accuracy for identifying named entities associated with heart disease risk factor events. The system's primary weakness was due to inaccuracies when characterizing the attributes of some events. For example, determining the relative time of an event with respect to the record date, whether an event is attributable to the patient's history or the patient's family history, and differentiating between current and prior smoking status. We believe these inaccuracies were due in large part to the lack of an effective approach for integrating context into our event detection model. To address these inaccuracies, we explore the addition of a distributional semantic model for characterizing contextual evidence of heart disease risk factor events. Using this semantic model, we raise our initial 2014 i2b2 Challenges in Natural Language of Clinical data F1 score of 0.838 to 0.890 and increased precision by 10.3% without use of any lexicons that might bias our results. Force-directed graph of diabetes mellitus and cad concepts extracted with distributional semantic model.