Using local lexicalized rules to identify heart disease risk factors in clinical notes.

Using local lexicalized rules to identify heart disease risk factors in clinical notes.
复制标题

DOI:
10.1016/j.jbi.2015.06.013
复制
发表时间:
2015-12
影响因子:
4.5
通讯作者:
Nenadic G
Nenadic G
中科院分区:
医学3区
文献类型:
--
作者:
Karystianis G;Dehghan A;Kovacevic A;Keane JA;Nenadic G

文献摘要

被引文献

相似文献

心脏病是全球主要的死亡原因,人类中有很大一部分人患有心脏病。许多危险因素已被认为是导致这种疾病的原因,包括肥胖、冠状动脉疾病(CAD)、高血压、高脂血症、糖尿病、吸烟和早产儿CAD家族史。本文描述和评估了一种从糖尿病临床笔记中提取这些危险因素的提及的方法,这是i2b2/UTHealth 2014临床数据自然语言处理挑战赛的任务之一。该方法是知识驱动的,系统实现了本地词汇化规则(基于笔记中观察到的句法模式),并结合了手动构建的表征该领域的词典。这项任务的一部分也是检测患者存在风险因素的时间间隔。该系统被应用于514个隐形纸币的评估集,获得了88%的微平均F分数(准确率为86%,召回率为90%)。虽然在识别冠心病家族史、用药和一些相关疾病因素(如高血压、糖尿病、高脂血症)方面取得了相当好的结果,但识别冠心病特异性指标被证明是更具挑战性的(F-Score为74%)。总体而言,结果令人鼓舞,并表明可以使用自动文本挖掘方法来处理临床笔记,以识别风险因素并大规模监测心脏病的进展,为临床和流行病学研究提供必要的数据。
Heart disease is the leading cause of death globally and a significant part of the human population lives with it. A number of risk factors have been recognised as contributing to the disease, including obesity, coronary artery disease (CAD), hypertension, hyperlipidemia, diabetes, smoking, and family history of premature CAD. This paper describes and evaluates a methodology to extract mentions of such risk factors from diabetic clinical notes, which was a task of the i2b2/UTHealth 2014 Challenge in Natural Language Processing for Clinical Data. The methodology is knowledge-driven and the system implements local lexicalised rules (based on syntactical patterns observed in notes) combined with manually constructed dictionaries that characterize the domain. A part of the task was also to detect the time interval in which the risk factors were present in a patient. The system was applied to an evaluation set of 514 unseen notes and achieved a micro-average F-score of 88% (with 86% precision and 90% recall). While the identification of CAD family history, medication and some of the related disease factors (e.g. hypertension, diabetes, hyperlipidemia) showed quite good results, the identification of CAD-specific indicators proved to be more challenging (F-score of 74%). Overall, the results are encouraging and suggested that automated text mining methods can be used to process clinical notes to identify risk factors and monitor progression of heart disease on a large-scale, providing necessary data for clinical and epidemiological studies.