Building Causal Graphs from Medical Literature and Electronic Medical Records

Building Causal Graphs from Medical Literature and Electronic Medical Records
复制标题

DOI:
10.1609/aaai.v33i01.33011102
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Galia Nordon;G. Koren;V. Shalev;B. Kimelfeld;Uri Shalit;Kira Radinsky
Galia Nordon;G. Koren;V. Shalev;B. Kimelfeld;Uri Shalit;Kira Radinsky
中科院分区:
其他
文献类型:
--
作者:
Galia Nordon;G. Koren;V. Shalev;B. Kimelfeld;Uri Shalit;Kira Radinsky

文献摘要

被引文献

相似文献

大型医疗数据库,如电子病历(EMR)数据,被认为是有前途的知识发现来源。对这些存储库的有效分析通常需要彻底理解数据中的依赖关系。例如,如果忽略患者年龄,则可能错误地得出白内障和高血压之间的因果关系。这种混杂变量通常通过因果图来识别,其中变量通过因果关系连接。目前自动构建这种图的方法是基于对医学文献的文本分析;然而,结果通常是低精度的大图。有一些统计方法可以从观测数据中构建因果图,但它们不太适合处理大量协变量,这就是EMR数据的情况。因此,混淆变量通常由医学领域专家通过手动、昂贵且耗时的过程来识别。我们提出了一种新的方法,用于自动构建医疗条件之间的因果图。第一部分是一种新的基于图形的方法,以更好地捕捉医学文献中隐含的因果关系,特别是在存在多个因果因素的情况下。然而,即使使用了这些先进的文本分析方法,文本数据仍然包含许多薄弱或不确定的因果关系。因此,我们基于超过150万患者的EMR库为这些术语构建了第二个图。我们联合收割机将这两个图表结合起来,只留下同时具有基于医学文本和观察证据的边缘。我们研究了几种策略来执行我们的方法,并使用医学专家比较了所得图形的精度。我们的研究结果表明,与最先进的方法相比,我们的任何方法的精度都有显着提高。
Large repositories of medical data, such as Electronic Medical Record (EMR) data, are recognized as promising sources for knowledge discovery. Effective analysis of such repositories often necessitate a thorough understanding of dependencies in the data. For example, if the patient age is ignored, then one might wrongly conclude a causal relationship between cataract and hypertension. Such confounding variables are often identified by causal graphs, where variables are connected by causal relationships. Current approaches to automatically building such graphs are based on text analysis over medical literature; yet, the result is typically a large graph of low precision. There are statistical methods for constructing causal graphs from observational data, but they are less suitable for dealing with a large number of covariates, which is the case in EMR data. Consequently, confounding variables are often identified by medical domain experts via a manual, expensive, and time-consuming process. We present a novel approach for automatically constructing causal graphs between medical conditions. The first part is a novel graph-based method to better capture causal relationships implied by medical literature, especially in the presence of multiple causal factors. Yet even after using these advanced text-analysis methods, the text data still contains many weak or uncertain causal connections. Therefore, we construct a second graph for these terms based on an EMR repository of over 1.5M patients. We combine the two graphs, leaving only edges that have both medical-text-based and observational evidence. We examine several strategies to carry out our approach, and compare the precision of the resulting graphs using medical experts. Our results show a significant improvement in the precision of any of our methods compared to the state of the art.