Patient Medical History Representation, Extraction, and Inference from EHR Data
Patient Medical History Representation, Extraction, and Inference from EHR Data
批准号:
9115724
负责人:
Cui Tao
金额:
$33.52万
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2019-08-31
关键词:
AddressAdoptedAftercareArchivesAutomated AnnotationBig DataChronic DiseaseClinicalClinical DataClinical ResearchColorectal CancerCommunicationComplexComputer softwareDataData CollectionData ReportingData SetData SourcesDatabasesDecision Support SystemsDetectionDiabetes MellitusDiseaseDisease ProgressionElectronic Health RecordEvaluationEventGoalsGoldHarvestHumanInstitutesMapsMeasuresMedical HistoryMedical RecordsModelingNatural Language ProcessingNatureOntologyPatient CarePatientsPerformanceRecording of previous eventsRegistriesReportingResolutionSemanticsStructureSystemTestingTimeTranslational ResearchWorkapplication programming interfacebaseclinical practicecohortcolon cancer patientsdata modelingdata structureinformation modelinnovationnovel strategiesopen sourcepersonalized medicinetooltrend analysis
中文摘要
描述(由申请人提供):开发从电子健康记录(EHR)中自动获取临床事件的时间约束的工具的意义怎么估计都不为过。对EHR数据中时间方面的有效分析可以促进一系列临床和转化性研究,如疾病进展研究、决策支持系统和个性化医学。
我们面临的一大挑战是自动解开嵌入高度多样化的大规模EHR数据中的临床事件的时间约束并将其线性化。时态数据建模、规范化、提取和推理的障碍阻碍了EHR数据源的有效利用,以进行事件历史评估和趋势分析:(1)现有的联邦支持的EHR数据规范化工具还没有关注非结构化数据的时间方面;(2)现有的时间模型只关注具有绝对时间的结构化数据,缺乏支持的推理系统,或者只提供特定于应用的局部解决方案,无法被复杂的EHR数据采用;(3)现有的时态信息提取方法要么难以适用于EHR数据,不具有可扩展性,要么只提供特定于应用的局部解决方案。
这个拟议的项目填补了目前本体、自然语言处理(NLP)和基于EHR的临床研究之间的空白,用于时间数据的表示、规范化、提取和推理。我们建议开发新的方法来自动表示、规范化和推理海量、多样化和异构性的EHR数据,并为进一步分析准备集成的数据。我们将在我们的Timer(时态信息建模、提取和推理)框架上构建新的推理和提取能力,以提供一个端到端的、开源的、符合标准的软件包。计时器将建立在强大的前期工作由我们的团队。我们将在我们的CNTRO(临床叙事时间关系本体)中开发新的功能,用于在复杂的EHR数据中语义定义时间域和表示时间数据。在新开发的CNTRO语义的基础上,我们将实现时间关系推理能力,以自动规范化时间表达式,计算和推断时间关系,并解决歧义。我们将利用现有的自然语言处理工具,并在这些工具的基础上开发新的提取方法,以填补当前自然语言处理方法和基于本体的推理方法之间的差距。我们将采用SHARPn EHR数据标准化管道和cTAKES来从临床叙述中提取和标准化临床事件提及。我们将探索一种创新的时态关系提取和事件共指方法,并使其与定时器框架一起工作。我们将使用来自两个机构的糖尿病(DM)和结直肠癌(CRC)患者队列来评估该系统。每个组成部分将首先进行单独测试,然后对整个框架进行评估。将报告精确度、召回率和f测量等结果。
英文摘要
DESCRIPTION (provided by applicant): The significance of developing tools for automatically harvesting temporal constraints of clinical events from Electronic Health Records (EHR) cannot be overestimated. Efficient analysis of the temporal aspects in EHR data could boost an array of clinical and translational research such as disease progression studies, decision support systems, and personalized medicine.
One big challenge we are facing is to automatically untangle and linearize the temporal constraints of clinical events embedded in highly diverse large-scale EHR data. Barriers to temporal data modeling, normalization, extraction, and reasoning have precluded the efficient use of EHR data sources for event history evaluation and trending analysis: (1) The current federally-supported EHR data normalization tools do not focus on the time aspect of unstructured data yet; (2) Existing time models focus only on structured data with absolute time, lack of supporting reasoning systems, or only offer application-specific partial solutions which cannot be adopted by the complex EHR data; (3) Current temporal information extraction approaches are either difficult to be adopted to EHR data, not scalable, or only offers application-specific partial solution.
This proposed project fills in the current gaps among ontologies, Natural Language Processing (NLP), and EHR-based clinical research for temporal data representation, normalization, extractions, and reasoning. We propose to develop novel approaches for automatic temporal data representation, normalization and reasoning for large, diverse, and heterogeneous EHR data and prepare the integrated data for further analysis. We will build new reasoning and extraction capacities on our TIMER (Temporal Information Modeling, Extracting, and Reasoning) framework to provide an end-to-end, open-source, standard-conforming software package. TIMER will be built on strong prior work by our team. We will develop new features in our CNTRO (Clinical Narrative Temporal Relation Ontology) for semantically defining the time domain and representing temporal data in complex EHR data. On top of the new developed CNTRO semantics, we will implement temporal relation reasoning capacities to automatically normalize temporal expressions, compute and infer temporal relations, and resolve ambiguities. We will leverage existing NLP tools and work on top of these tools to develop new extraction approaches to fill in the current gaps between NLP approaches and ontology-based reasoning approaches. We will adapt the SHARPn EHR data normalization pipeline and cTAKES for extracting and normalizing clinical event mentions from clinical narratives. We will explore an innovative approach for temporal relation extraction and event coreference, and make it work with the TIMER framework. We will evaluate the system using Diabetes Mellitus (DM) and colorectal cancer (CRC) patient cohorts from two insititutions. Each component will be tested separately first followed by an evaluation of the whole framework. Results such as precision, recall, and f-measure will be reported.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Metadata applications on informed content to facilitate biorepository data regulation and sharing
-
批准号:9360131
-
项目类别:
-
资助金额:$45.67万
-
财政年份:2016
-
负责人:Cui Tao
-
依托单位:
Patient Medical History Representation, Extraction, and Inference from EHR Data
-
批准号:8760594
-
项目类别:
-
资助金额:$39.83万
-
财政年份:2014
-
负责人:Cui Tao
-
依托单位:
海外基金