Evaluating the coverage of controlled health data terminologies: Report on the results of the NLM/AHCPR large scale vocabulary test

Evaluating the coverage of controlled health data terminologies: Report on the results of the NLM/AHCPR large scale vocabulary test
复制标题

DOI:
10.1136/jamia.1997.0040484
复制
发表时间:
1997-11-01
影响因子:
6.4
通讯作者:
Cheh, ML
Cheh, ML
中科院分区:
管理学2区
文献类型:
--
作者:
Humphreys, BL;McCray, AT;Cheh, ML

文献摘要

被引文献

相似文献

目的:通过使用互联网和 UMLS 知识源、词汇程序和服务器进行分布式国家实验,确定现有机器可读健康术语的组合在多大程度上涵盖了健康信息系统综合受控词汇所需的概念和术语。方法:使用专门设计的基于 Web 的 UMLS 知识源服务器界面,参与者检索了 1996 年 UMLS 元同义词库中的 30 多个词汇和三个计划添加的词汇,以确定概念是否存在他们希望存在或不存在受控术语。对于提交的每个术语,界面都会呈现候选精确匹配或一组潜在的近似匹配,参与者从中选择最密切相关的概念。该界面捕获了参与者提交的术语的概况,以及对于每个搜索术语,有关参与者选择的概念(如果有)的信息。术语信息被加载到 NLM 的数据库中以供审查和分析,并且也可供参与者下载。一组主题专家审查了记录,以识别参与者错过的比赛并纠正关系中任何明显的错误。 SNOMED International 和 Read Codes 的编辑获得了已审核术语的随机样本,其中未找到确切含义匹配的术语,以识别遗漏的精确匹配或与输入术语同义的概念的任何有效组合。 1997 UMLS Metathesaurus 被用于语义类型和词汇源分析,因为它包含了三个计划添加的大部分。结果:63 位参与者总共提交了 41,127 个术语,代表 32,679 个规范化字符串。超过 80% 的提交条款是针对与患者病情相关的患者记录部分。经审查,所有提交的术语中,58% 的术语与测试中受控词汇的含义完全匹配,41% 的术语具有相关概念,1% 的术语未找到。在 28% 的术语中,其含义比受控词汇表中的概念更窄,其中 86% 与更广泛的概念共享词汇项,但有额外的修改;准确含义匹配的百分比因专业而异,从 45% 到 71% 不等。 29 个不同的词汇表包含 23,837 个术语(最多 12,707 个离散概念)中的一些的含义,并且含义完全匹配。根据初步数据和分析,个别词汇包含
Objective: To determine the extent to which a combination of existing machine-readable health terminologies cover the concepts and terms needed for a comprehensive controlled vocabulary for health information systems by carrying out a distributed national experiment using the Internet and the UMLS Knowledge Sources,lexical programs, and server.Methods: Using a specially designed Web-based interface to the UMLS Knowledge Source Server, participants searched the more than 30 vocabularies in the 1996 UMLS Metathesaurus and three planned additions to determine if concepts for which they desired controlled terminology were present or absent. For each term submitted, the interface presented a candidate exact match or a set of potential approximate matches from which the participant selected the most closely related concept. The interface captured a profile of the terms submitted by the participant and for each term searched, information about the concept (if any) selected by the participant. The term information was loaded into a database at NLM for review and analysis and was also available to be downloaded by the participant. A team of subject experts reviewed records to identify matches missed by participants and to correct any obvious errors in relationships. The editors of SNOMED International and the Read Codes were given a random sample of reviewed terms for which exact meaning matches were not found to identify exact matches that were missed or any valid combinations of concepts that were synonymous to input terms. The 1997 UMLS Metathesaurus was used in the semantic type and vocabulary source analysis because it included most of the three planned additions.Results: Sixty-three participants submitted a total of 41,127 terms, which represented 32,679 normalized strings. Mure than 80% of the terms submitted were wanted for parts of the patient record related to the patient's condition. Following review, 58% of all submitted terms had exact meaning matches in the controlled vocabularies in the test, 41% had related concepts, and 1% were not found. Of the 28% of the terms which were narrower in meaning than a concept in the controlled vocabularies, 86% shared lexical items with the broader concept, but had additional modification; The percentage of exact meaning matches varied by specialty from 45% to 71%. Twenty-nine different vocabularies contained meanings for some of the 23,837 terms (a maximum of 12,707 discrete concepts) with exact meaning matches. Based on preliminary data and analysis, individual vocabularies contained