A Methodology for Reliable Risk Assessment with Error-prone Electronic Medical Records Using Optimal Design of Experiments Concepts
A Methodology for Reliable Risk Assessment with Error-prone Electronic Medical Records Using Optimal Design of Experiments Concepts
批准号:
1436574
负责人:
Daniel Apley
金额:
$40.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-08-31
中文摘要
大量的医疗保健资源致力于编辑电子病历(EMR)数据库,该数据库在患者群体中越来越集成和丰富,并且提供了经由统计分析来识别疾病风险因素以预测作为各种因素(例如,临床和人口统计学)。不幸的是,疾病事件数据可能具有高的错误编码错误率,这是由于雇用了具有有限培训的文书人员来输入他们的代码。例如,在一个心脏检查患者的EMR数据库中,在审查了记录为心脏骤停事件的病例的随机样本后,发现错误率为75%。为了考虑这些错误并避免开发不可靠的风险评估模型,医生必须进行图表审查,以验证病例样本并确定事件是否为真实事件。然而,由于医生的时间成本很高,病历审查的数量有限。本研究的目的是开发一种方法,明智和有效地选择验证的情况下,最大的信息内容,这将允许可靠的疾病风险评估,即使是高度易错的EMR数据。对社会健康和福祉的预期益处是巨大的,因为这项研究将使大型EMR数据库的巨大未开发潜力得到更充分的利用,以发现新的疾病风险因素。预计该研究还可以扩展到其他大数据应用领域,以便从大量质量可疑的数据中提取可靠的信息。大型电子病历(EMR)数据库通过拟合统计模型提供了发展临床假设和识别疾病风险关联的可能性,所述统计模型预测患者发展特定状况的可能性作为各种预测因子的函数变量(例如,该患者的临床、表型和人口统计数据)。虽然预测变量数据通常被可靠地记录,但由于ICD-9疾病错误编码,事件数据可能具有高错误率。为了避免开发不可靠的风险评估模型,以前的研究使用随机验证抽样来估计错误概率,以纠正Logistic回归模型拟合整个数据的偏差,这是效率低下且不可靠的高错误率。与此相反,本研究将开发一个验证抽样和可靠的风险评估(VSRRA)的方法,明智地设计一个验证样本。知识基础是观察到的VSRRA和传统的实验设计(DOE)之间的类比,从而验证响应一个错误的情况下,VSRRA对应于进行一个实验运行在DOE。根据这种类比,本研究将开发(i)适用于医疗风险研究中常用的广义线性模型的基于Fisher信息矩阵的模型参数和贝叶斯对应物(如后验和前验参数协方差矩阵)的VSRRA设计标准;(ii)用于选择验证样本以优化设计标准的启发式和更精确的混合算法;(iii)多级、顺序版本的VSRRA采样策略,其基于在新情况被验证时沿着学习的信息来改进设计;以及(iv)确定是否以及如何在最终模型拟合中沿着可靠地包括未验证数据的完整集合以及已验证数据的方法。数据分析的一个基本原则是,精心设计的实验研究比观察研究产生更可靠的统计结论。同样,预计基于DOE的VSRRA方法将允许更可靠的疾病风险评估和假设生成。
英文摘要
Enormous healthcare resources are devoted to compiling electronic medical record (EMR) databases that are increasingly integrated and rich in patient population and that offer potential for identifying disease risk factors via statistical analyses to predict the disease risk as a function of various factors (e.g., clinical and demographic) for that patient. Unfortunately, the disease event data may have high miscoding error rates, due to the fact that clerical personnel with limited training are employed to enter their codes. For example, in one EMR database of patients with cardiac workup, after reviewing a random sample of cases recorded as sudden cardiac arrest events, the error rate was found to be 75 percent. In order to take such errors into account and avoid developing unreliable risk assessment models, it is imperative that a doctor perform chart reviews to validate a sample of cases and determine whether the events were true events. However, the number of chart reviews is limited due to the high cost of doctors' time. The objective of this research is to develop a methodology for judiciously and efficiently selecting validation cases for maximum information content, which will allow reliable disease risk assessment even with highly error-prone EMR data. The anticipated benefits to the health and well-being of society are substantial, as this research will allow the enormous untapped potential of large EMR databases to be more fully utilized for discovering new disease risk factors. It is also anticipated that this research can be extended to other big-data application domains for extracting reliable information from large quantities of data that are of questionable quality.Large electronic medical record (EMR) databases offer potential for developing clinical hypotheses and identifying disease risk associations by fitting statistical models that predict the likelihood that a patient develops a particular condition as a function of various predictor variables (e.g., clinical, phenotypical, and demographic data) for that patient. Although the predictor variable data are often recorded reliably, the event data may have high error rates due to ICD-9 disease miscoding. To avoid developing unreliable risk assessment models, previous research used random validation sampling to estimate error probabilities for correcting biases in logistic regression models fit to the entire data, which is both inefficient and unreliable with high error rates. In contrast, this research will develop a validation sampling and reliable risk assessment (VSRRA) methodology for judiciously designing a validation sample. The intellectual underpinning is the observed analogy between VSRRA and traditional design of experiments (DOE), whereby validating the response for one error-prone case in VSRRA corresponds to conducting one experimental run in DOE. In light of this analogy, this research will develop (i) suitable VSRRA design criteria based on the Fisher information matrix for the model parameters and Bayesian counterparts such as posterior and preposterior parameter covariance matrices, applicable to a broad class of generalized linear models commonly used in medical risk studies; (ii) heuristic and more exact hybrid algorithms for selecting the validation sample to optimize the design criteria; (iii) multistage, sequential versions of the VSRRA sampling strategies that refine the designs based on information that is learned along the way, as new cases are validated; and (iv) methods that determine whether and how the full set of unvalidated data can be reliably included, along with the validated data, in the final model fitting. A fundamental tenet of data analysis is that carefully designed experimental studies produce far more reliable statistical conclusions than observational studies. Likewise, it is anticipated that the DOE-based VSRRA methodology will allow far more reliable disease risk assessment and hypotheses generation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Model-Based Multidisciplinary Dynamic Decisions in Design
-
批准号:1537641
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2015
-
负责人:Daniel Apley
-
依托单位:
Collaborative Research: Leveraging Noncontact Dimensional Metrology to Understand Complex Part-to-Part Variation
-
批准号:1265709
-
项目类别:Standard Grant
-
资助金额:$18.26万
-
财政年份:2013
-
负责人:Daniel Apley
-
依托单位:
Enhancing Identifiability of Computer Simulation Models via Design for Calibration
-
批准号:1233403
-
项目类别:Standard Grant
-
资助金额:$32.0万
-
财政年份:2012
-
负责人:Daniel Apley
-
依托单位:
Collaborative Research: Blind Discovery of Variation Sources for Visualization by Multidisciplinary Teams
-
批准号:0826081
-
项目类别:Standard Grant
-
资助金额:$18.99万
-
财政年份:2008
-
负责人:Daniel Apley
-
依托单位:
A Bayesian Treatment of Uncertainty in Simulation-Based Methods for Enhancing Process and Product Robustness
-
批准号:0758557
-
项目类别:Standard Grant
-
资助金额:$32.0万
-
财政年份:2008
-
负责人:Daniel Apley
-
依托单位:
CAREER: A Methodology to Systematically Characterize and Diagnose Manufacturing Variation with In-Process Measurement Data
-
批准号:0354824
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Daniel Apley
-
依托单位:
CAREER: A Methodology to Systematically Characterize and Diagnose Manufacturing Variation with In-Process Measurement Data
-
批准号:0093580
-
项目类别:Continuing Grant
-
资助金额:$37.5万
-
财政年份:2001
-
负责人:Daniel Apley
-
依托单位:
海外基金