Frequent Causal Pattern Mining: A Computationally Efficient Framework For Estimating Bias-Corrected Effects.

Frequent Causal Pattern Mining: A Computationally Efficient Framework For Estimating Bias-Corrected Effects.
复制标题

DOI:
10.1109/bigdata47090.2019.9005977
复制
发表时间:
2019-12
期刊:
Proceedings : ... IEEE International Conference on Big Data. IEEE International Conference on Big Data
影响因子:
--
通讯作者:
Simon G
Simon G
中科院分区:
其他
文献类型:
--
作者:
Yadav P;Caraballo PJ;Steinbach M;Kumar V;Castro MR;Simon G

文献摘要

参考文献

相似文献

我们的老龄化人口越来越多地同时患有多种慢性病,需要对这些疾病进行综合治疗。为疾病的组合集合寻找最优药物集合是一个组合模式探索问题。关联规则挖掘是解决这类问题的一种流行工具,但医疗保健需要发现因果模式而不是关联模式,这使得关联规则挖掘变得不合适。为了解决这个问题,我们提出了一个基于Rubin-Neyman因果模型的新框架,用于从观测数据中提取因果规则,纠正了一些常见的偏差。具体地说,给定一组干预措施和一组定义亚群体(例如,疾病)的项目,我们希望找到其中存在有效干预组合的所有子群体,并且在每个这样的子群体中,我们希望找到所有干预组合,使得从该组合中丢弃任何干预将降低治疗的有效性。我们框架的一个关键方面是封闭干预集合的概念,它将量化单个干预的影响的概念扩展到一组同时进行的干预。封闭干预集还允许使用一种修剪策略,该策略严格地比Apriori算法使用的传统修剪策略更有效。为了实现我们的想法,我们介绍并比较了从观测数据估计因果效应的五种方法,并在合成数据上严格评估它们,以数学证明(在可能的情况下)它们为什么有效。我们还在梅奥诊所152000名患者的电子健康记录数据上评估了我们的因果规则挖掘框架,结果表明我们提取的模式足够丰富,足以解释医学文献中关于一类胆固醇药物对II型糖尿病(T2 DM)的影响的有争议的发现。
Our aging population increasingly suffers from multiple chronic diseases simultaneously, necessitating the comprehensive treatment of these conditions. Finding the optimal set of drugs for a combinatorial set of diseases is a combinatorial pattern exploration problem. Association rule mining is a popular tool for such problems, but the requirement of health care for finding causal, rather than associative, patterns renders association rule mining unsuitable. To address this issue, we propose a novel framework based on the Rubin-Neyman causal model for extracting causal rules from observational data, correcting for a number of common biases. Specifically, given a set of interventions and a set of items that define subpopulations (e.g., diseases), we wish to find all subpopulations in which effective intervention combinations exist and in each such subpopulation, we wish to find all intervention combinations such that dropping any intervention from this combination will reduce the efficacy of the treatment. A key aspect of our framework is the concept of closed intervention sets which extend the concept of quantifying the effect of a single intervention to a set of concurrent interventions. Closed intervention sets also allow for a pruning strategy that is strictly more efficient than the traditional pruning strategy used by the Apriori algorithm. To implement our ideas, we introduce and compare five methods of estimating causal effect from observational data and rigorously evaluate them on synthetic data to mathematically prove (when possible) why they work. We also evaluated our causal rule mining framework on the Electronic Health Records (EHR) data of a large cohort of 152000 patients from Mayo Clinic and showed that the patterns we extracted are sufficiently rich to explain the controversial findings in the medical literature regarding the effect of a class of cholesterol drugs on Type-II Diabetes Mellitus (T2DM).
DOI: 10.1016/j.jbi.2011.07.001
发表时间: 2011-12
影响因子: 4.5
作者:
Kleinberg, Samantha;Hripcsak, George
通讯作者: Hripcsak, George
DOI: 10.1023/a:1009891813863
发表时间: 2000-07-01
影响因子: 4.8
作者:
Silverstein, C;Brin, S;Ullman, J
通讯作者: Ullman, J
DOI: 10.1080/00273171.2011.540480
发表时间: 2011
影响因子: 3.8
作者:
Austin PC
通讯作者: Austin PC
DOI: 10.1097/ccm.0000000000002949
发表时间: 2018-04
影响因子: 8.8
作者:
Pruinelli L;Westra BL;Yadav P;Hoff A;Steinbach M;Kumar V;Delaney CW;Simon G
通讯作者: Simon G
DOI: 10.1023/a:1009787925236
发表时间: 1997-01-01
影响因子: 4.8
作者:
Cooper, GF
通讯作者: Cooper, GF