Maximizing clinical cohort size using free text queries.

Maximizing clinical cohort size using free text queries.
复制标题

使用自由文本查询最大化临床队列规模。

DOI:
10.1016/j.compbiomed.2015.01.008
复制
发表时间:
2015
影响因子:
7.7
通讯作者:
Zeng-Treitler,Qing
Zeng-Treitler,Qing
中科院分区:
工程技术2区
文献类型:
--
作者:
Gundlapalli,AdiV;Redd,Doug;Gibson,BryanSmith;Carter,Marjorie;Korhonen,Chris;Nebeker,Jonathan;Samore,MatthewH;Zeng-Treitler,Qing

文献摘要

相似文献

背景队列识别在人群健康管理和研究中都具有重要意义。在这个项目中,我们试图评估文本查询用于队列识别的使用情况。具体地说,我们试图确定将非结构化数据查询添加到结构化查询中以进行患者队列识别时的增量价值。方法评估了三种队列识别任务:同时服用银杏叶和华法林(银杏/华法林)的个体识别、超重的个体识别和未控制的糖尿病(UCD)个体识别。我们评估了将非结构化数据查询添加到结构化数据查询中时队列大小的增加。结果对于银杏/华法林患者,文本查询将队列大小从9个增加到28,924个,而仅通过查询药房数据来确定队列。对于与体重相关的任务,文本搜索比通过查询VITALS表确定的队列增加了5%-29%的队列。对于UCD任务,与通过查询实验室结果或ICD代码确定的队列相比,文本查询使队列大小增加了2%-43%。文本搜索对银杏/华法林的阳性预测值为52%,权重队列的阳性预测值为19-94%,UCD的阳性预测值为44%。讨论本项目展示了自由文本查询在从大数据集中识别患者队列的价值和局限性。纳入和排除标准在患者群体中的临床范围和流行率影响这一方法的实用性和效果。
BackgroundCohort identification is important in both population health management and research. In this project we sought to assess the use of text queries for cohort identification. Specifically we sought to determine the incremental value of unstructured data queries when added to structured queries for the purpose of patient cohort identification.MethodsThree cohort identification tasks were evaluated: identification of individuals taking gingko biloba and warfarin simultaneously (Gingko/Warfarin), individuals who were overweight, and individuals with uncontrolled diabetes (UCD). We assessed the increase in cohort size when unstructured data queries were added to structured data queries. The positive predictive value of unstructured data queries was assessed by manual chart review of a random sample of 500 patients.ResultsFor Gingko/Warfarin, text query increased the cohort size from 9 to 28,924 over the cohort identified by query of pharmacy data only. For the weight-related tasks, text search increased the cohort by 5–29% compared to the cohort identified by query of the vitals table. For the UCD task, text query increased the cohort size by 2–43% compared to the cohort identified by query of laboratory results or ICD codes. The positive predictive values for text searches were 52% for Gingko/Warfarin, 19–94% for the weight cohort and 44% for UCD.DiscussionThis project demonstrates the value and limitation of free text queries in patient cohort identification from large data sets. The clinical domain and prevalence of the inclusion and exclusion criteria in the patient population influence the utility and yield of this approach.