Integrative data science approaches for rare disease discovery in health records
Integrative data science approaches for rare disease discovery in health records
批准号:
9884791
负责人:
Vikas Rao Pejaver
金额:
$9.21万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-03-04 至 2022-02-28
关键词:
AdultAffectAmericanAwardBasic ScienceBehavioralBioinformaticsClinicalClinical DataClinical MedicineClinical ResearchComputing MethodologiesConsensusDataData ScienceData SetDetectionDiagnosisDiagnosticDiagnostics ResearchDiseaseEconomic BurdenElectronic Health RecordEnvironmentFacultyFamilyGenesGeneticGenomicsGenotypeGoalsHealthcareHealthcare SystemsIndividualInformaticsKnowledgeMachine LearningManuscriptsMarkov ChainsMedicalMedical GeneticsMental disordersMentorsMentorshipMethodsMiningModelingMolecularNamesNatural Language ProcessingNatural Language Processing pipelineOntologyOutcomePacific NorthwestPatient RecruitmentsPatientsPatternPhasePhenotypePopulationPositioning AttributePrevalencePrincipal InvestigatorRare DiseasesRecording of previous eventsResearchResearch PersonnelStandardizationSymptomsSystemTestingTimeTrainingUniversitiesValidationVocabularyWashingtonWorkaccurate diagnosisbasebiomedical data sciencebiomedical informaticscareercausal variantclinical data warehouseclinical decision-makingcohortdiagnostic accuracydisease phenotypeearly onset disorderexome sequencinggene discoverygenomic datahealth care deliveryhealth datahealth recordimprovedmembermultimodal datanovelopen sourcepatient health informationphenotypic dataprototypepsychologicrare conditionrare genetic disorderrecruitskillssoftware developmentsupport toolstooltrait
中文摘要
点击翻译按钮获取中文摘要
英文摘要
ABSTRACT: There are nearly 7,000 diseases that have a prevalence of only one in 2,000 individuals or less.
Yet, such rare diseases are estimated to collectively affect over 300 million people worldwide, representing a
significant healthcare concern. Although rare diseases have predominantly genetic origins, nearly half of them
do not manifest symptoms until adulthood and frequently confound discovery and diagnosis. Even in the case
of early onset disorders, the sheer number of possible diagnoses can often overwhelm clinicians. As a result,
rare diseases are often diagnosed with delay, misdiagnosed or even remain undiagnosed, not only disrupting
patient lives but also hindering progress on our understanding of such diseases. Data science methods that
mine large-scale retrospective health record data for phenotypic information will aid in timely and accurate
diagnoses of rare diseases, especially when combined with additional data types, thus, having significant real-
world impact. This proposal will integrate electronic health record (EHR) data sets with publicly available
vocabularies and ontologies, and genomic data for the improved identification and characterization of patients
with rare diseases, using approaches from machine learning, natural language processing (NLP) and basic
bioinformatics. The work has three specific aims and will be carried out in two phases. During the mentored
phase, the principal investigator (PI) will develop data-driven methods to extract standardized concepts related
to rare diseases from clinical notes and infer the occurrence of each disease (Aim 1). He will also develop data
science approaches to compare and contrast longitudinal patterns associated with patients' journeys through
the healthcare system when seeking a diagnosis for a rare disease, and aid in clinical decision-making by
leveraging these patterns (Aim 2). During the independent phase (Aim 3), computational methods will be
developed for the integrated modeling and analysis of genotypic (from Aim 3) and phenotypic information (from
Aims 1 and 2). Cohorts to be sequenced will cover diseases for which causal genes or disease definitions are
unclear (discovery), as well as those for which these are well known (validation). This work will be carried out
under the mentorship of four faculty members with complementary expertise in biomedical informatics, data
science, NLP, and rare disease genomics at the University of Washington, the largest medical system in the
Pacific Northwest (four million EHRs), world-renowned researchers in medical genetics, and a robust data
science environment. In addition, under the direction of the mentoring team, the PI will complete advanced
coursework, receive training in translational bioinformatics and clinical research informatics, submit
manuscripts, and seek an independent research position. This proposal will yield preliminary results for
subsequent studies on data-driven phenotyping and enable the realization of the PI's career goals by providing
him with the necessary training to build on his machine learning and basic bioinformatics expertise to transition
into an independent investigator in biomedical data science.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Integrative data science approaches for rare disease discovery in health records
-
批准号:10626148
-
项目类别:
-
资助金额:$23.65万
-
财政年份:2022
-
负责人:Vikas Rao Pejaver
-
依托单位:
Integrative data science approaches for rare disease discovery in health records
-
批准号:10541283
-
项目类别:
-
资助金额:$23.59万
-
财政年份:2022
-
负责人:Vikas Rao Pejaver
-
依托单位:
海外基金