Automated linking of free-text complaints to reason-for-visit categories and International Classification of Diseases diagnoses in emergency department patient record databases

Automated linking of free-text complaints to reason-for-visit categories and International Classification of Diseases diagnoses in emergency department patient record databases
复制标题

DOI:
10.1016/s0196-0644(03)00748-0
复制
发表时间:
2004-03-01
影响因子:
6.2
通讯作者:
La, M
La, M
中科院分区:
医学1区
文献类型:
--
作者:
Day, FC;Schriger, DL;La, M

文献摘要

被引文献

相似文献

研究目的:使用国际疾病分类系统来描述急诊科(艾德)病例组合存在缺陷。因此,我们开发了计算机算法,可以识别单词,单词片段和单词模式的组合,将自由文本投诉字段与20个访问原因类别联系起来。我们研究的可行性和可靠性,将这些原因的访问类别艾德患者visitdatabases.Methods:我们分析了一个数据库(包含投诉和国际疾病分类诊断为I年的访问一个单一的艾德)使用3步的过程(创建初始条款,最大限度地提高敏感性,最大限度地提高特异性),以定义20个原因的访问类别的纳入和排除条款。为了评估访视原因分配算法的可靠性,我们在第二个数据库上重复了最后2个步骤,该数据库由21个ED的访视样本组成。对于每个数据库,我们确定了与每个就诊原因类别相关的投诉患病率,以及所有患者和按年龄分层的患者的国际疾病分类第九版诊断分布。结果:20个访视原因类别涵盖了数据库I中所有患者的77%(平均年龄33.5岁)和数据库2中所有患者的67%(平均年龄38.9岁)。对于数据库I和2,按年龄范围划分的20个访视原因类别采集的访视百分比分别为0 - 2岁(84%和76%)、3 - 10岁(82%和74%)、11 - 65岁(76%和68%)和66岁或以上(69%和60%)。数据库之间与每个就诊原因类别相关的所有投诉的比例基本相似。每个投诉领域,是链接到每个访问的原因类别包括至少一个术语,它涉及到类别标题,和最频繁分配的诊断,在每个访问的原因类别是那些人们会期望与访问的原因类别controls.Conclusion:的方法,其中自由文本投诉字段被解析为访问的原因类别是可行的,合理可靠的;最后确定的数据库I访问原因类别纳入/排除术语列表仅需要适度的修改就可以在数据库2中良好地工作。此处使用的访视原因类别定义较广,以最大限度地提高其捕获的访视比例;定义较窄的访视原因类别在不同数据库中使用时,需要对其入选/排除术语列表进行更广泛的修订。一个前瞻性的,基于就诊原因的艾德分类系统可能有几个有用的应用(包括症状监测),虽然内容效度分析将是必要的,以调查这一假设。
Study objective: The use of the International Classification of Diseases system to describe emergency department (ED) case mix has disadvantages. We therefore developed computer algorithms that recognize a combination of words, word fragments, and word patterns to link free-text complaint fields to 20 reason-for-visit categories. We examine the feasibility and reliability of applying these reason-for-visit categories to ED patient-visit databases.Methods: We analyzed a database (containing complaints and International Classification of Diseases diagnoses for I year's visits to a single ED) using a 3-step process (create initial terms, maximize sensitivity, maximize specificity) to define inclusion and exclusion terms for 20 reason-for-visit categories. To assess the reliability of the reason-for-visit assignment algorithm, we repeated the final 2 steps on a second database, composed of visits sampled from 21 EDs. For each database, we determined the prevalence of complaints that link to each reason-for-visit category and the distributions of International Classification of Diseases, Ninth Revision diagnoses that resulted for all patients and patients stratified by age.Results: The 20 reason-for-visit categories capture 77% of all patients in database I (mean age 33.5 years) and 67% of all patients in database 2 (mean age 38.9 years). The percentage of visits captured by the 20 reason-for-visit categories, by age range, for databases I and 2 are (respectively) 0 to 2 years (84% and 76%), 3 to 10 years (82% and 74%), 11 to 65 years (76% and 68%), and 66 years or older (69% and 60%). The proportions of all complaints that link to each reason-for-visit category are largely similar between databases. Every complaint field that is linked to each reason-for-visit category includes at least I term that relates it to the category title, and the most frequently assigned diagnoses in each reason-for-visit category are those that one would expect to be associated with the reason-for-visit category complaints.Conclusion: The method by which free-text complaint fields are parsed into reason-for-visit categories is feasible and reasonably reliable; the finalized database I reason-for-visit category inclusion/exclusion terms lists required only modest changes to work well in database 2. The reason-for-visit categories used here are broadly defined to maximize the proportion of visits that they capture; more narrowly defined reason-for-visit categories will require more extensive revision of their inclusion/exclusion terms lists when used in different databases. A prospective, reason-for-visit-based ED classification system could have several useful applications (including syndromic surveillance), although content validity analysis will be necessary to investigate this hypothesis.