Statistical Methods for Incorporating Machine Learning Tools in Inference and Large-Scale Surveillance using Electronic Medical Records Data
Statistical Methods for Incorporating Machine Learning Tools in Inference and Large-Scale Surveillance using Electronic Medical Records Data
批准号:
10463566
负责人:
Marco Carone
金额:
$48.63万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-07-18 至 2024-06-30
关键词:
AlgorithmsBenefits and RisksCharacteristicsClinical DataComplexComputer softwareComputerized Medical RecordConfidence IntervalsCountryDataData CollectionData SetDetectionDimensionsE-learningEarly DiagnosisEffectivenessElectronic Health RecordEnsureEpidemiologyEstimation TechniquesEventFrequenciesGoalsHealthHealthcareHeterogeneityInformation SystemsInfrastructureInterventionKnowledgeLearningLinkMachine LearningMeasuresMethodologyMethodsModernizationMonitorOutcomePatient-Focused OutcomesPatientsPharmacodynamicsPopulationPopulation SurveillanceProceduresPublic HealthRecordsReproducibilityResearchResearch MethodologyResearch PersonnelSafetySentinelSignal TransductionStandardizationStatistical MethodsStreamStructural ModelsSubgroupSurveillance ProgramSystemTechniquesTestingTimeTreatment outcomeUpdatebaseclinical carecomparative treatmentdata streamsflexibilityhigh dimensionalityimplementation barriersimprovedinterestmachine learning methodnational surveillancenovelopen sourcepatient populationpatient subsetsrisk minimizationsoftware developmentsurveillance datatooltreatment effectuser-friendly
中文摘要
摘要
临床护理信息系统的现代化和标准化正在创建大型网络
链接的电子健康记录(EHR),可捕获关键治疗并选择患者结果
全国数以百万计的患者。从这些系统中产生的观测数据
提供一个无与伦比的机会来了解现有和新的治疗方法的有效性,
并监测在广大患者中使用干预措施时可能出现的潜在安全问题
人口。然而,观察性临床数据的暴露是由许多因素和
因此,需要进行积极的调整,以消除尽可能多的混杂偏见,以便
对选定的暴露进行归因。机器学习领域提供了一个强大的
用于执行灵活、彻底的混杂调整的数据驱动方法的集合,但
当这些技术用于以下用途时,执行可靠的统计推断尤其具有挑战性
分析策略的一部分。我们建议通过开发和开发可重复使用的研究方法来推进研究方法
说明利用机器学习方法的灵活性的新的有针对性的学习工具
使用大规模电子病历数据检测和表征健康影响信号。
具体地说,我们将首先开发用于进行高效、统计有效和健壮的推理的技术
使用最先进的机器学习工具实现治疗效果。我们还将发展在线学习
在流式传输EHR数据的上下文中进行此类推断的技术。方法学上的进步将
使我们能够制定一个正式、严谨和实用的框架,以进行持续、有效的
以及对安全终端的可靠监控。最后,我们将开发统计方法,用于
纳入先验信息--包括人口统计学、流行病学或药效学
知识,例如--改善健康结果时对健康影响的估计和推断
感兴趣的情况很少见,因此统计问题很困难,就像在安全监测中经常发生的那样。
拟议研究的最终目标是使生物医学研究人员和公共卫生
监管机构仔细监测和保护公众的健康,允许他们更有效地
并更可靠地检测可能包含在人口规模的EHR中的关键健康影响信号
数据。
英文摘要
SUMMARY
The modernization and standardization of clinical care information systems is creating large networks of
linked electronic health records (EHR) that capture key treatments and select patient outcomes for
millions of patients throughout the country. The observational data emerging from these systems
provide an unparalleled opportunity to learn about the effectiveness of existing and novel treatments,
and to monitor potential safety issues that may arise when interventions are used in broad patient
populations. However, observational clinical data have exposures that are driven by many factors and
therefore aggressive adjustment is needed to remove as much confounding bias as possible in order to
make attribution regarding select exposures. The field of machine learning provides a powerful
collection of data-driven approaches for performing flexible, thorough confounding adjustment, but
performing reliable statistical inference is particularly challenging when these techniques are used as
part of the analytic strategy. We propose to advance reproducible research methods by developing and
illustrating novel targeted learning tools that leverage the flexibility of machine learning methods to
detect and characterize health effect signals using large-scale EHR data.
Specifically, we will first develop techniques for making efficient, statistically valid and robust inference
for treatment effects using state-of-the-art machine learning tools. We will also develop online learning
techniques to make such inference in the context of streaming EHR data. Methodological advances will
enable us to formulate a formal, rigorous and practical framework for conducting continuous, effective
and reliable surveillance for safety endpoints. Finally, we will develop statistical approaches for
incorporating prior information -- including demographic, epidemiologic or pharmacodynamic
knowledge, for example -- to improve health effect estimation and inference when the health outcome
of interest is rare and the statistical problem is thus difficult, as often occurs in safety surveillance.
The ultimate goal of the proposed research is to enable biomedical researchers and public health
regulators to carefully monitor and protect the health of the public by allowing them to more effectively
and more reliably detect critical health effect signals that may be contained in population-scale EHR
data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Methods for Incorporating Machine Learning Tools in Inference and Large-Scale Surveillance using Electronic Medical Records Data
-
批准号:9816009
-
项目类别:
-
资助金额:$50.28万
-
财政年份:2019
-
负责人:Marco Carone
-
依托单位:
Statistical Methods for Incorporating Machine Learning Tools in Inference and Large-Scale Surveillance using Electronic Medical Records Data
-
批准号:9979940
-
项目类别:
-
资助金额:$48.6万
-
财政年份:2019
-
负责人:Marco Carone
-
依托单位:
Statistical Methods for Incorporating Machine Learning Tools in Inference and Large-Scale Surveillance using Electronic Medical Records Data
-
批准号:10645177
-
项目类别:
-
资助金额:$48.65万
-
财政年份:2019
-
负责人:Marco Carone
-
依托单位:
Statistical Methods for Incorporating Machine Learning Tools in Inference and Large-Scale Surveillance using Electronic Medical Records Data
-
批准号:10206237
-
项目类别:
-
资助金额:$48.62万
-
财政年份:2019
-
负责人:Marco Carone
-
依托单位:
海外基金