A data science framework for transforming electronic health records into real-world evidence
A data science framework for transforming electronic health records into real-world evidence
批准号:
10664706
负责人:
Vivek A Rudrapatna
金额:
$8.9万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-03 至 2025-07-31
关键词:
3-DimensionalAccelerationAlgorithmsBayesian NetworkBenchmarkingBiometryCaliforniaChronicClassificationClinic VisitsClinical DataClinical ResearchClinical TrialsComplexDataData ReportingData ScienceData SetDedicationsDiseaseDrug ApprovalE-learningEffectivenessElderlyElectronic Health RecordEligibility DeterminationEndoscopyEquityExclusionFood and Drug Administration Drug ApprovalFutureGoalsHealthcareImmune System DiseasesJointsLearningMachine LearningMalignant NeoplasmsMasksMeasurementMeasuresMentorsMethodsModelingNatural Language ProcessingNatureNew Drug ApprovalsPatient RepresentativePatientsPatternPerformancePharmaceutical PreparationsPopulationPredispositionPregnancyPublishingRaceRandomized, Controlled TrialsRecording of previous eventsResearch SubjectsSample SizeSan FranciscoSelection for TreatmentsSemanticsSeveritiesSourceStructureSubgroupSymptomsTestingTextTimeTrainingTreatment EffectivenessUlcerative ColitisUncertaintyUniversitiesWorkalgorithm trainingcareercareer developmentcohortcomputerized toolscostdata harmonizationdata integrationelectronic structureheterogenous dataimprovedin silicoinnovationinsightlearning strategymeetingsoutcome predictionpatient health informationpatient subsetsrandomized trialreconstructionsupport toolstooltreatment effect
中文摘要
点击翻译按钮获取中文摘要
英文摘要
PROJECT SUMMARY
Randomized controlled trials (RCTs) are the gold-standard in clinical research but are subject to many
limitations including high costs, limited generalizability, and small sample sizes in patient subgroups. By
contrast, electronic health records (EHRs) are widely available and contain information on large and
representative patient cohorts. However, because they capture the uncontrolled observations of many
clinicians, they are highly susceptible to bias. The recent availability of the raw data from RCTs has created a
unique opportunity to integrate them with that from EHRs, and to innovate methods that exploit the distinct
advantages of each dataset.
We propose to identify the zone of overlap between these data and build bridges in data representations.
These bridges could enable us to better emulate randomized trials using EHR data and measure the same
effects seen in the trials. Consequently, it would allow us to study subgroups that were excluded from the
pivotal trials associated with new drug approvals by the FDA.
We will test these ideas out in the context of Ulcerative colitis (UC) and scale to others in future work. We
have obtained access to the raw data from 12 RCTs in UC (N=6,226). These data contain timed and structured
measurements of disease activity including the Mayo score, a composite score of patient symptoms and
endoscopic severity. We have also obtained access to the EHR data of 3,270 UC patients treated at the
University of California San Francisco. These data contain similar data as RCTs but largely in an unstructured
form. In addition, these assessments tend to be incomplete relative to trials due to costs and invasiveness of
some tests. We will address this problem of unharmonized and incomplete EHR data in three aims.
In Aim 1, we will harmonize the RCT data into an analysis-ready format. We will also develop text
classification tools to transform free-texted EHR data into Mayo subscores, and validate these tools against
data from a second center. In Aim 2, we will integrate the RCT and EHR data, train algorithms to impute RCT-
based representations of the patient state from partial measurements made in EHRs, and test them under
conditions typifying real-world data capture. In Aim 3, we will use these algorithms to harmonize EHR data,
validate them as a tool to recover the same effects as RCTs, and study new patient subgroups.
The applicant will carry out these aims and train in biostatistics, natural language processing, machine
learning, and overall career development. With the help of his mentors, he will launch a career dedicated to
developing and disseminating methods for learning from complex clinical data, and in so doing, promote a
future of better healthcare for all patients.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Algorithmic Identification of Treatment-Emergent Adverse Events From Clinical Notes Using Large Language Models: A Pilot Study in Inflammatory Bowel Disease.
使用大型语言模型从临床记录中算法识别治疗中出现的不良事件:炎症性肠病的初步研究。
DOI:
10.1002/cpt.3226
发表时间:
2024
期刊:
Clinical pharmacology and therapeutics
影响因子:
6.7
作者:
[Silverman,AnnaL, Sushil,Madhumita, Bhasuran,Balu, Ludwig,Dana, Buchanan,James, Racz,Rebecca, Parakala,Mahalakshmi, El-Kamary,Samer, Ahima,Ohenewaa, Belov,Artur, Choi,Lauren, Billings,Monisha, Li,Yan, Habal,Nadia, Liu,Qi, Tiwari,Jawahar, B]
通讯作者:
B
DOI:
10.1093/ibd/izad041
发表时间:
2023-03
期刊:
Inflammatory bowel diseases
影响因子:
4.9
作者:
[F. Odufalu;Justin L Sewell;Vivek A. Rudrapatna;M. Somsouk;U. Mahadevan]
通讯作者:
F. Odufalu;Justin L Sewell;Vivek A. Rudrapatna;M. Somsouk;U. Mahadevan
海外基金