Research data warehouse best practices: catalyzing national data sharing through informatics innovation.

Research data warehouse best practices: catalyzing national data sharing through informatics innovation.
复制标题

研究数据仓库最佳实践:通过信息学创新促进国家数据共享。

DOI:
10.1093/jamia/ocac024
复制
发表时间:
2022
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Lenert,LeslieA
Lenert,LeslieA
中科院分区:
--
文献类型:
--
作者:
Murphy,ShawnN;Visweswaran,Shyam;Becich,MichaelJ;Campion,ThomasR;Knosp,BoydM;Melton-Meaux,GenevieveB;Lenert,LeslieA

文献摘要

相似文献

研究患者数据存储库(RPDR)已成为传统临床和翻译科学奖(CTSA)计划的基本基础设施,并日益成为范围广泛的研究联盟1-4和学习卫生系统网络的基础设施。几乎每个拥有CTSA或临床转化研究(CTR)计划的机构(发现在国家卫生研究院资助额较低的州)都会为附属研究人员提供RPDR。这些资料库旨在根据它们所服务的患者群体进行医疗保健研究。在机构内部,区域PDR对于一系列研究活动是有价值的。它们被用来识别临床试验招募的患者,使用保护隐私的方法来搜索和提取符合试验条件的患者的特定队列。它们有助于开发和验证可计算的表型,这些表型对于准确和可重复地识别患者队列越来越重要。RPDR为人口健康研究提供识别的患者数据,并支持越来越多的人工智能工作来预测患者结果。此外,临床研究通常可以使用来自RPDR的数据来模拟。在该机构之外,来自多个机构的已识别数据集的聚合与隐私保护哈希码相关联,提供了一个前所未有的机会,可以进行人口健康研究,执行比较有效性分析,并在大量和多样化的人口中应用人工智能方法。总体而言,RPDR在加快翻译研究方面的好处可能很大。例如,在哈佛大学,2006年,每年9400至1.36亿美元的研究资金与使用RPDR的数据有关。5根据机构的长处和短处,《区域发展报告》所载数据因机构而异;本期发表的论文反映了这种差异(见表1)。最常见的情况是,数据从本地电子健康记录(EHR)和其他临床信息系统获取,这些系统在临床护理期间捕获信息。数据包括诊断、问题清单、程序、处方药、实验室检查和许多类型的自由文本报告。如表1所示,一些RPDR使用自然语言处理方法增强了研究人员可用的编码数据。RPDR越来越多地包含其他类型的数据,包括来自生物库样本的基因组数据、临床试验数据、调查数据,如患者报告的结果,以及来自医疗保险索赔的数据。医疗和电子设备已成为重要的数据来源,包括在医院内获得的数据,如成像和重症监护病房监测,以及医院外从可穿戴设备和家庭测量获得的数据。调查问卷中关于健康的社会决定因素的数据、州和国家死亡指数中的死亡数据以及地理编码的环境数据,包括潜在的毒素和与天气有关的数据,也越来越多地可用。这期《Jamia》特刊介绍了当前RDPR的一些研究、方法、应用和最佳实践。该问题包括这一领域广泛的积极研究,包括11篇研究和应用论文6-16篇和4份案例报告。17-20
Research Patient Data Repositories (RPDRs) have become essential infrastructure for traditional Clinical and Translational Science Award (CTSA) programs and increasingly for a wide range of research consortia 1–4 and learning health system networks. Almost every institution with a CTSA or Clinical Translational Research (CTR) program (found in states with lower amounts of National Institutes of Health funding) hosts an RPDR for the benefit of affiliated researchers. These repositories aim to enable healthcare research based upon the patient populations they serve. Within the institution, RPDRs are valuable for a range of research activities. They are used to identify patients for clinical trial recruitment using privacy-preserving methods to search and extract specific cohorts of trial-eligible patients. They aid in the development and validation of computable phenotypes that are increasingly important for identifying patient cohorts accurately and in a reproducible fashion. RPDRs provide deidentified patient data for population health research and support a growing body of artificial intelligence work for predicting patient outcomes. Further, clinical studies can often be simulated using data from an RPDR. Beyond the institution, aggregates of deidentified datasets from multiple institutions linked with privacy-preserving hash codes provide an unprecedented opportunity to conduct population health research, perform comparative effectiveness analyses, and apply artificial intelligence methods over large and diverse populations. Overall, the benefits of the RPDR for accelerating translational research can be large. For example, at Harvard, in 2006, between $94 and $136 million in annual research funding was linked to use of data from the RPDR. 5 The data contained within the RPDR vary across institutions, based on institutional strengths and weaknesses; the papers published in this issue reflect that variability (see Table 1). Most commonly, data are acquired from local electronic health record (EHR) and other clinical information systems that captured information during clinical care. Data consist of diagnoses, problem lists, procedures, prescribed medications, laboratory exams, and many types of free-text reports. As shown in Table 1, some RPDRs enhance the coded data available to researchers using natural language processing methods. RPDRs increasingly contain additional types of data, including genomic data derived from biobanked samples, clinical trial data, survey data such as patient-reported outcomes, and data from health insurance claims. Medical and electronic devices have become important sources of data, including that obtained within the hospital, such as imaging and intensive care unit monitoring, and that obtained outside the hospital from wearables and home measurements. Data on social determinants of health from questionnaires, death data from state and national death indexes, and geocoded environmental data, including potential toxins and weatherrelated data, are also increasingly available. This special issue of JAMIA describes some of the current research, approaches, applications, and best practices for RDPRs. The issue includes a wide range of active research in this area and includes 11 research and applications papers 6–16 and 4 case reports. 17–20