Detecting Lung and Colorectal Cancer Recurrence Using Structured Clinical/Administrative Data to Enable Outcomes Research and Population Health Management.

Detecting Lung and Colorectal Cancer Recurrence Using Structured Clinical/Administrative Data to Enable Outcomes Research and Population Health Management.
复制标题

使用结构化临床/行政数据检测肺和大肠癌复发,以实现结果和人群健康管理。

DOI:
10.1097/mlr.0000000000000404
复制
发表时间:
2017-12
期刊:
影响因子:
3
通讯作者:
Ritzwoller D
Ritzwoller D
中科院分区:
医学3区
文献类型:
--
作者:
Hassett MJ;Uno H;Cronin AM;Carroll NM;Hornbrook MC;Ritzwoller D

文献摘要

被引文献

相似文献

复发性癌症是常见的,昂贵的,致命的,但我们对社区人群知之甚少。电子健康记录(EHR)和肿瘤登记处包含大量关于社区患者的数据,但通常缺乏复发状态。使用结构化数据来检测复发的现有算法具有局限性。我们开发了算法来检测I-III期肺癌和结直肠癌确定性治疗后复发的存在和时间,使用两个数据源,这些数据源包含与金标准复发状态相关的广泛可用的结构化数据类型(索赔或EHR遭遇):与癌症护理结果研究和监测研究相关的医疗保险索赔,以及与注册数据相关的癌症研究网络虚拟数据仓库。12个潜在的复发指标被用于为每个数据源中的每种癌症开发单独的模型。检测模型使ROC曲线下面积(AUC)最大化;计时模型使平均绝对误差最小化。通过癌症类型/数据源比较算法,并与现有的二进制检测规则进行对比。检测模型AUC(>0.92)超过了现有的预测规则。时间模型产生的绝对预测误差相对于随访时间较小(<15%)。相似的协变量被纳入所有检测和定时算法中,尽管癌症类型和数据集的差异挑战了为所有场景创建一个通用算法的努力。使用大数据进行有效和可靠的复发检测是可行的。这些工具将使广泛的,新颖的研究质量,有效性和结果的肺癌和结直肠癌患者和那些谁开发复发。
Recurrent cancer is common, costly, and lethal, yet we know little about it in community-based populations. Electronic health records (EHR) and tumor registries contain vast amounts of data regarding community-based patients, but usually lack recurrence status. Existing algorithms that use structured data to detect recurrence have limitations. We developed algorithms to detect the presence and timing of recurrence after definitive therapy for stages I-III lung and colorectal cancer using two data sources that contain a widely available type of structured data (claims or EHR encounters) linked to gold standard recurrence status: Medicare claims linked to the Cancer Care Outcomes Research and Surveillance study, and the Cancer Research Network Virtual Data Warehouse linked to registry data. Twelve potential indicators of recurrence were used to develop separate models for each cancer in each data-source. Detection models maximized area under the ROC curve (AUC); timing models minimized average absolute error. Algorithms were compared by cancer type/data-source, and contrasted with an existing binary detection rule. Detection model AUC’s (>0.92) exceeded existing prediction rules. Timing models yielded absolute prediction errors that were small relative to follow-up time (<15%). Similar covariates were included in all detection and timing algorithms, though differences by cancer-type and dataset challenged efforts to create one common algorithm for all scenarios. Valid and reliable detection of recurrence using big data is feasible. These tools will enable extensive, novel research on quality, effectiveness, and outcomes for lung and colorectal cancer patients and those who develop recurrence.