Semi-structured Information Retrieval in Clinical Text for Cohort Identification
Semi-structured Information Retrieval in Clinical Text for Cohort Identification
批准号:
8811565
负责人:
HONGFANG LIU
金额:
$46.07万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-20 至 2019-07-31
关键词:
AccountingAddressAdoptedAdoptionAsthmaClinicClinicalCollectionCommunitiesComputerized Medical RecordComputersDataDictionaryDiseaseElectronic Health RecordEpidemiologistEpidemiologyEvaluationEventEvidence Based MedicineEvolutionGoalsHealthInformation RetrievalInformation Retrieval SystemsInstitutionInterest GroupInvestigationJudgmentLanguageLearningMachine LearningMeasuresMedicalMedical RecordsMethodologyMethodsMetricModelingModificationMorphologic artifactsNamesNatural Language ProcessingOutcomePatient RecruitmentsPatientsPerformancePharmaceutical PreparationsPhasePhysiciansProcessPublishingQualifyingRecordsResearchResearch PersonnelResourcesRestRetrievalSamplingSemanticsSiteSmokeSourceSpecific qualifier valueStructureSystemTechniquesTestingTextValidationWeightWorkWritingasthmatic patientbasecohortimprovedindexingnovelopen sourcepublic health relevancesyntaxtext searchingtool
中文摘要
描述(由申请人提供):自然语言处理(NLP)技术已经显示出从电子健康记录(EHR)的自由文本中提取数据的前景,但研究一致地发现,该技术不容易在应用程序设置中推广。不幸的是,将NLP应用于实际用例的大部分注意力仍然停留在单一的、定义良好的应用程序设置的范例上,因此对看不见的用例的概括性仍然没有得到解决。我们建议通过采用信息检索(IR)的角度来明确说明未见的应用程序设置,以患者级别队列识别为目标。为此,我们引入了分层语言模型,这是一个支持重用NLP生成的构件的IR框架。我们的长期目标是通过提供强大的、以用户为中心的工具来加速对患者健康和疾病的调查,这些工具对于处理、检索和利用EHR的免费文本是必要的。这项建议的主要目标是从梅奥诊所和OHSU的临床文本中准确地检索特别的、现实的队列,为患者水平的IR建立方法、资源和评估。我们假设,队列识别可以通过一个新的IR框架以一种可推广的方式解决:分层语言模型。我们将通过四个具体目标来检验这一假设。在目标1中,我们将使医疗NLP构件在我们的分层语言IR框架中可搜索。这包括存储和索引NLP构件,以及使用统计语言模型基于文本及其关联的NLP构件检索文档。在目标2中,我们处理特别队列识别的实际设置,转移到患者级(而不是文档级)IR。为了准确地处理可能将合格证据分散在多个文档中的患者队列,我们将开发和实现患者级别的检索模型,该模型考虑了事件的跨文档关系和时间组合。在目标3中,我们将使用来自两个站点的EHR数据构建并行的IR测试集合;由多个站点编写的不同队列查询集
人们对各种临床或流行病学目的的了解;以及评估哪些患者与两个站点上的哪些查询相关。最后,在目标4中,我们在特别队列识别任务中精炼和评估患者级别的分层语言IR,跨用户、查询、优化指标和机构进行比较。我们将与先前存在的技术进行额外的外部比较,例如,针对来自电子医疗记录和基因组网络的队列。拟议工作的预期结果是:(I)临床医生和流行病学家可使用的开放源码队列识别工具,该工具原则上利用NLP人工制品进行看不见的查询;ii)队列识别的并行测试集合,包括两个机构内部文档集合、不同的测试主题和用户生成的文本查询,以及与每个查询相关的患者级别的判断;以及(Iii)通过检索患者队列任务验证医学NLP的可重用性。
英文摘要
DESCRIPTION (provided by applicant): Natural Language Processing (NLP) techniques have shown promise for extracting data from the free text of electronic health records (EHRs), but studies have consistently found that techniques do not readily generalize across application settings. Unfortunately, most of the focus in applying NLP to real use cases has remained on a paradigm of single, well-defined application settings, so that generalizability to unseen use cases remains implicitly unaddressed. We propose to explicitly account for unseen application settings by adopting an information retrieval (IR) perspective with the objective of patient-level cohort identification. To do so, we introduce layered language models, an IR framework that enables the reuse of NLP-produced artifacts. Our long term goal is to accelerate investigations of patient health and disease by providing robust, user- centric tools that are necessary to process, retrieve, and utilize the free text of EHRs. The main goal of this proposal is to accurately retrieve ad hoc, realistic cohorts from clinical text at Mayo Clinic and OHSU, establishing methods, resources, and evaluation for patient-level IR. We hypothesize that cohort identification can be addressed in a generalizable fashion by a new IR framework: layered language models. We will test this hypothesis through four specific aims. In Aim 1, we will make medical NLP artifacts searchable in our layered language IR framework. This involves storing and indexing the NLP artifacts, as well as using statistical language models to retrieve documents based on text and its associated NLP artifacts. In Aim 2, we deal with the practical setting of ad hoc cohort identification, moving to patient-level (rather than document-level) IR. To accurately handle patient cohorts in which qualifying evidence may be spread over multiple documents, we will develop and implement patient-level retrieval models that account for cross- document relational and temporal combinations of events. In Aim 3, we will construct parallel IR test collections using EHR data from two sites; a diverse set of cohort queries written by multiple
people toward various clinical or epidemiological ends; and assessments of which patients are relevant to which queries at both sites. Finally, in Aim 4, we refine and evaluate patient-level layered language IR on the ad hoc cohort identification task, making comparisons across the users, queries, optimization metrics, and institutions. We will draw additional extrinsic comparisons with pre-existing techniques, e.g., for cohorts from the Electronic Medical Records and Genonmics network. The expected outcomes of the proposed work are: (i) An open-source cohort identification tool, usable by clinicians and epidemiologists, that makes principled use of NLP artifacts for unseen queries; ii) A parallel test collection for cohort identification, includig two intra-institutional document collections, diverse test topics and user-produced text queries, and patient-level judgments of relevance to each query; and (iii) Validation of the reusability of medical NLP via the task of retrieving patient cohorts.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Learning Precision Medicine for Rare Diseases Empowered by Knowledge-driven Data Mining
-
批准号:10732934
-
项目类别:
-
资助金额:$72.37万
-
财政年份:2023
-
负责人:HONGFANG LIU
-
依托单位:
The Data, Evaluation, and Coordination Center (DECC) for Connecting Underrepresented Populations to Clinical Trials (CUSP2CT)
-
批准号:10597291
-
项目类别:
-
资助金额:$55.44万
-
财政年份:2022
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10202598
-
项目类别:
-
资助金额:$63.08万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10001498
-
项目类别:
-
资助金额:$64.37万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:9251814
-
项目类别:
-
资助金额:$30.0万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10471838
-
项目类别:
-
资助金额:$64.37万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Semi-structured Information Retrieval in Clinical Text for Cohort Identification
-
批准号:8928647
-
项目类别:
-
资助金额:$37.63万
-
财政年份:2014
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:9033918
-
项目类别:
-
资助金额:$56.28万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8640959
-
项目类别:
-
资助金额:$58.01万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8920720
-
项目类别:
-
资助金额:$16.0万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8505753
-
项目类别:
-
资助金额:$63.07万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8826771
-
项目类别:
-
资助金额:$57.16万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
-
批准号:8448471
-
项目类别:
-
资助金额:$61.4万
-
财政年份:2009
-
负责人:HONGFANG LIU
-
依托单位:
Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
-
批准号:7654995
-
项目类别:
-
资助金额:$60.87万
-
财政年份:2009
-
负责人:HONGFANG LIU
-
依托单位:
海外基金