A semantic lexicon for medical language processing

A semantic lexicon for medical language processing
复制标题

DOI:
10.1136/jamia.1999.0060205
复制
发表时间:
1999-05-01
影响因子:
6.4
通讯作者:
Johnson, SB
Johnson, SB
中科院分区:
管理学2区
文献类型:
--
作者:
Johnson, SB

文献摘要

被引文献

相似文献

目的:构建一个提供词语和短语语义信息的资源,以便于计算机处理医学叙述。设计:将专家词典中的词位(词语和词语短语)与美国国家医学图书馆1997年出版的《统一医学语言系统(UMLS)元词库》中的字符串进行匹配。这就产生了一个“语义词典”,其中每个词位与一个或多个句法类型相关联,每个句法类型可以有一个或多个语义类型。然后,使用语义词典为放电摘要语料库(603,306个句子)中出现的词汇分配语义类型。考察了具有多种语义类型的词项,以确定是否可以根据在放电摘要中的使用情况来确定是否可以消除某些类型的语义。一个整合程序被用来为每个反映不同语义的词位寻找对比语境。结果:将专家词典与元主题词表进行匹配,得到一个包含75,711种词汇形式的语义词典,其中22,805种(30.1%)有两种或两种以上的语义类型与专业词典相匹配,其中有27,633种不同的词汇类型,其中13,322种至少有一种语义类型。这表明,专家词典对放电摘要的句法信息覆盖率约为79%,对语义信息的覆盖率约为38%。在语料库中具有语义类型的词汇中,3474个(12.6%)具有两种或两种以上类型。当语义偏好规则应用于语义词典时,具有多种语义类型的词条数量减少到423个(1.5%)。在出库摘要中,具有多种语义类型的词汇出现率从9.41%下降到1.46%。结论:可以使用自动方法从现有的UMLS来源构建语义词典。这种语义信息可以帮助自然语言处理程序分析医疗叙事,前提是将具有多种语义类型的词位保持在最少。语义偏好规则可用于选择适合临床报告的语义类型。还需要进一步的工作来增加语义词典的覆盖面,并在选择语义时利用上下文信息。
Objective: Construction of a resource that provides semantic information about words and phrases to facilitate the computer processing of medical narrative.Design: Lexemes (words and word phrases) in the Specialist Lexicon were matched against strings in the 1997 Metathesaurus of the Unified Medical Language System (UMLS) developed by the National Library of Medicine. This yielded a "semantic lexicon," in which each lexeme is associated with one or more syntactic types, each of which can have one or more semantic types. The semantic lexicon was then used to assign semantic types to lexemes occurring in a corpus of discharge summaries (603,306 sentences). Lexical items with multiple semantic types were examined to determine whether some of the types could be eliminated, on the basis of usage in discharge summaries. A concordance program was used to find contrasting contexts for each lexeme that would reflect different semantic senses. Based on this evidence, semantic preference rules were developed to reduce the number of lexemes with multiple semantic types.Results: Matching the Specialist Lexicon against the Metathesaurus produced a semantic lexicon with 75,711 lexical forms, 22,805 (30.1 percent) of which had two or more semantic types Matching the Specialist Lexicon against one year's worth of discharge summaries identified 27,633 distinct lexical forms, 13,322 of which had at least one semantic type. This suggests that the Specialist Lexicon has about 79 percent coverage for syntactic information and 38 percent coverage for semantic information for discharge summaries. Of those lexemes in the corpus that had semantic types, 3,474 (12.6 percent) had two or more types. When semantic preference rules were applied to the semantic lexicon, the number of entries with multiple semantic types was reduced to 423 (1.5 percent). In the discharge summaries, occurrences of lexemes with multiple semantic types were reduced from 9.41 to 1.46 percent.Conclusion: Automatic methods can be used to construct a semantic lexicon from existing UMLS sources. This semantic information can aid natural language processing programs that analyze medical narrative, provided that lexemes with multiple semantic types are kept to a minimum. Semantic preference rules can be used to select semantic types that are appropriate to clinical reports. Further work is needed to increase the coverage of the semantic lexicon and to exploit contextual information when selecting semantic senses.