Generate Analysis-Ready Data for Real-world Evidence: Tutorial for Harnessing Electronic Health Records With Advanced Informatic Technologies.

Generate Analysis-Ready Data for Real-world Evidence: Tutorial for Harnessing Electronic Health Records With Advanced Informatic Technologies.
复制标题

DOI:
10.2196/45662
复制
发表时间:
2023-05-25
影响因子:
7.4
通讯作者:
--
中科院分区:
医学2区
文献类型:
--
作者:

文献摘要

相似文献

尽管随机对照试验 (RCT) 是确定药物治疗有效性和安全性的黄金标准,但根据真实世界数据生成的真实世界证据 (RWE) 在批准后监测中至关重要,并且正在推广用于实验疗法的监管过程。现实世界数据的一个新兴来源是电子健康记录(EHR),其中包含结构化(例如诊断代码)和非结构化(例如临床记录和图像)形式的患者护理详细信息。尽管电子病历中可用数据的粒度很细,但可靠评估治疗与临床结果之间关系所需的关键变量很难提取。为了解决这一基本挑战并加速 EHR 在 RWE 中的可靠使用,我们引入了一个集成的数据管理和建模管道,该管道由 4 个模块组成,这些模块利用自然语言处理、计算表型和噪声数据因果建模技术的最新进展。模块 1 包含数据协调技术。我们使用自然语言处理从 RCT 设计文档中识别临床变量,并通过描述匹配和知识网络将提取的变量映射到 EHR 特征。然后,模块 2 使用先进的表型分析算法开发队列构建技术,以识别患有感兴趣疾病的患者并定义治疗组。模块 3 介绍了变量管理的方法,包括一系列现有工具,用于从不同来源(例如,编码、自由文本和医学成像)和各种类型的终点(例如,死亡、二进制、时间和数字)提取基线变量。最后,模块 4 介绍了验证和稳健的建模方法,我们提出了一种策略,为感兴趣的 EHR 变量创建黄金标准标签,以验证数据管理质量并为 RWE 执行后续因果建模。除了我们管道中提出的工作流程之外,我们还为 RWE 制定了报告指南,其中涵盖了促进透明报告和结果可重复性所需的信息。此外,我们的管道是高度数据驱动的,通过丰富的公开信息和知识来源增强研究数据。我们还通过重新审视早期结肠癌患者腹腔镜辅助结肠切除术与开腹结肠切除术的手术治疗研究组试验的临床结果模拟来展示我们的产品线并为相关工具的部署提供指导。我们还借鉴了有关 RCT 的 EHR 模拟的现有文献以及我们自己与麻省总医院 Brigham EHR 的研究。
Although randomized controlled trials (RCTs) are the gold standard for establishing the efficacy and safety of a medical treatment, real-world evidence (RWE) generated from real-world data has been vital in postapproval monitoring and is being promoted for the regulatory process of experimental therapies. An emerging source of real-world data is electronic health records (EHRs), which contain detailed information on patient care in both structured (eg, diagnosis codes) and unstructured (eg, clinical notes and images) forms. Despite the granularity of the data available in EHRs, the critical variables required to reliably assess the relationship between a treatment and clinical outcome are challenging to extract. To address this fundamental challenge and accelerate the reliable use of EHRs for RWE, we introduce an integrated data curation and modeling pipeline consisting of 4 modules that leverage recent advances in natural language processing, computational phenotyping, and causal modeling techniques with noisy data. Module 1 consists of techniques for data harmonization. We use natural language processing to recognize clinical variables from RCT design documents and map the extracted variables to EHR features with description matching and knowledge networks. Module 2 then develops techniques for cohort construction using advanced phenotyping algorithms to both identify patients with diseases of interest and define the treatment arms. Module 3 introduces methods for variable curation, including a list of existing tools to extract baseline variables from different sources (eg, codified, free text, and medical imaging) and end points of various types (eg, death, binary, temporal, and numerical). Finally, module 4 presents validation and robust modeling methods, and we propose a strategy to create gold-standard labels for EHR variables of interest to validate data curation quality and perform subsequent causal modeling for RWE. In addition to the workflow proposed in our pipeline, we also develop a reporting guideline for RWE that covers the necessary information to facilitate transparent reporting and reproducibility of results. Moreover, our pipeline is highly data driven, enhancing study data with a rich variety of publicly available information and knowledge sources. We also showcase our pipeline and provide guidance on the deployment of relevant tools by revisiting the emulation of the Clinical Outcomes of Surgical Therapy Study Group Trial on laparoscopy-assisted colectomy versus open colectomy in patients with early-stage colon cancer. We also draw on existing literature on EHR emulation of RCTs together with our own studies with the Mass General Brigham EHR.