A context-blocks model for identifying clinical relationships in patient records.

A context-blocks model for identifying clinical relationships in patient records.
复制标题

DOI:
10.1186/1471-2105-12-s3-s3
复制
发表时间:
2011-06-09
期刊:
影响因子:
3
通讯作者:
Lu Z
Lu Z
中科院分区:
生物学4区
文献类型:
--
作者:
Islamaj Doğan R;Névéol A;Lu Z

文献摘要

被引文献

相似文献

患者记录包含关于诊断解释、疾病进展、处方和/或治疗效果等方面的有价值的信息。临床重要概念的自动识别和患者记录中这些概念之间的关系的识别是医学信息学中许多重要应用的基础步骤,从护理质量到假设生成。在这项工作中,我们描述了一种方法,该方法有助于自动识别在医疗问题、治疗和测试之间定义的八个关系。与传统的词袋表示不同,在这项工作中,我们使用由概念在文本中的位置确定的五个不同上下文块的方案来表示一种关系。作为关系识别的初步步骤,为了提供端到端系统,我们还解决了医疗问题、治疗和测试的自动提取。我们的方法结合了概念识别统计模型的结果和条件随机场模型中简单的自然语言处理特征。来自第四届i2b2挑战赛的826份患者记录被用于培训和评估该系统。结果表明,对于精确的跨度概念检测,我们的概念识别系统的F度量达到了0.870。此外,关系的上下文块表征(F-MEASURE=0.775)在识别关系方面比词袋(F-MEASURE=0.402)更成功。最重要的是,使用自动提取的概念的端到端关系提取系统的性能(F-MEASure=0.704)与使用手动注释的概念(F-MEASURE=0.711)获得的性能相当,并且它们的差异在统计学上不显著。我们以自动化的方式从文本中提取重要的临床关系,从概念识别开始,以关系识别结束。上下文块表示方案的优点是正确地管理单词位置信息,这在识别某些关系时可能是关键的。我们的结果可以作为基准,与基于i2b2挑战数据开发的其他系统进行比较。最后,我们的系统可以作为医学信息学中其他发现任务的初步步骤。
Patient records contain valuable information regarding explanation of diagnosis, progression of disease, prescription and/or effectiveness of treatment, and more. Automatic recognition of clinically important concepts and the identification of relationships between those concepts in patient records are preliminary steps for many important applications in medical informatics, ranging from quality of care to hypothesis generation. In this work we describe an approach that facilitates the automatic recognition of eight relationships defined between medical problems, treatments and tests. Unlike the traditional bag-of-words representation, in this work, we represent a relationship with a scheme of five distinct context-blocks determined by the position of concepts in the text. As a preliminary step to relationship recognition, and in order to provide an end-to-end system, we also addressed the automatic extraction of medical problems, treatments and tests. Our approach combined the outcome of a statistical model for concept recognition and simple natural language processing features in a conditional random fields model. A set of 826 patient records from the 4th i2b2 challenge was used for training and evaluating the system. Results show that our concept recognition system achieved an F-measure of 0.870 for exact span concept detection. Moreover the context-block representation of relationships was more successful (F-Measure = 0.775) at identifying relationships than bag-of-words (F-Measure = 0.402). Most importantly, the performance of the end-to-end system of relationship extraction using automatically extracted concepts (F-Measure = 0.704) was comparable to that obtained using manually annotated concepts (F-Measure = 0.711), and their difference was not statistically significant. We extracted important clinical relationships from text in an automated manner, starting with concept recognition, and ending with relationship identification. The advantage of the context-blocks representation scheme was the correct management of word position information, which may be critical in identifying certain relationships. Our results may serve as benchmark for comparison to other systems developed on i2b2 challenge data. Finally, our system may serve as a preliminary step for other discovery tasks in medical informatics.