Unified Medical Language System term occurrences in clinical notes: a large-scale corpus analysis.

Unified Medical Language System term occurrences in clinical notes: a large-scale corpus analysis.
复制标题

DOI:
10.1136/amiajnl-2011-000744
复制
发表时间:
2012-06
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Shah NH
Shah NH
中科院分区:
其他
文献类型:
--
作者:
Wu ST;Liu H;Li D;Tao C;Musen MA;Chute CG;Shah NH

文献摘要

参考文献

被引文献

相似文献

为了在大型临床语料库中验证统一医学语言系统(UMLS)元词库术语字符串的经验实例,并说明哪些类型的术语特征可在数据源中推广。基于5100万份马约诊所临床笔记的文档语料库中的UMLS术语的出现率,本研究计算了关于术语的字符串属性、源术语、语义类型和句法类别的统计。还映射了2010年i2 b2/VA文本中出现的术语;根据基于梅奥的统计数据设计了8个示例过滤器,并应用于i2 b2/VA数据。对于语料库分析,马约语料库中的映射术语的数量可以忽略不计,其具有超过6个单词或55个字符。在UMLS中的源术语中,消费者健康词汇和医学临床术语系统化命名(SNOMED-CT)在马约临床笔记中的覆盖率最高,分别为106 426和94 788个唯一术语。  在UMLS的15个语义组中,7个语义组占马约数据中术语出现的92.08%。句法上,超过90%的匹配项是名词短语。对于跨机构分析,在i2 b2/VA数据上使用五个示例过滤器将实际词汇减少到UMLS大小的19.13%,并且匹配术语仅减少2%。语料库的统计数据在这里提出的建设词汇从UMLS是有益的。Metathesaurus术语的固有特征(格式良好,长度和语言)很容易在临床机构中推广,但术语频率应谨慎调整。映射术语的语义组可能会因机构而略有不同,但当转移到生物医学文献领域时,它们会有很大不同。
To characterise empirical instances of Unified Medical Language System (UMLS) Metathesaurus term strings in a large clinical corpus, and to illustrate what types of term characteristics are generalisable across data sources. Based on the occurrences of UMLS terms in a 51 million document corpus of Mayo Clinic clinical notes, this study computes statistics about the terms' string attributes, source terminologies, semantic types and syntactic categories. Term occurrences in 2010 i2b2/VA text were also mapped; eight example filters were designed from the Mayo-based statistics and applied to i2b2/VA data. For the corpus analysis, negligible numbers of mapped terms in the Mayo corpus had over six words or 55 characters. Of source terminologies in the UMLS, the Consumer Health Vocabulary and Systematized Nomenclature of Medicine—Clinical Terms (SNOMED-CT) had the best coverage in Mayo clinical notes at 106 426 and 94 788 unique terms, respectively. Of 15 semantic groups in the UMLS, seven groups accounted for 92.08% of term occurrences in Mayo data. Syntactically, over 90% of matched terms were in noun phrases. For the cross-institutional analysis, using five example filters on i2b2/VA data reduces the actual lexicon to 19.13% of the size of the UMLS and only sees a 2% reduction in matched terms. The corpus statistics presented here are instructive for building lexicons from the UMLS. Features intrinsic to Metathesaurus terms (well formedness, length and language) generalise easily across clinical institutions, but term frequencies should be adapted with caution. The semantic groups of mapped terms may differ slightly from institution to institution, but they differ greatly when moving to the biomedical literature domain.
DOI: 10.1126/science.1199644
发表时间: 2011-01-14
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Michel JB;Shen YK;Aiden AP;Veres A;Gray MK;Google Books Team;Pickett JP;Hoiberg D;Clancy D;Norvig P;Orwant J;Pinker S;Nowak MA;Aiden EL
通讯作者: Aiden EL
DOI: 10.1186/1471-2105-12-397
发表时间: 2011-10-12
期刊: BMC bioinformatics
影响因子: 3
作者:
Thompson P;McNaught J;Montemagni S;Calzolari N;del Gratta R;Lee V;Marchi S;Monachini M;Pezik P;Quochi V;Rupp CJ;Sasaki Y;Venturi G;Rebholz-Schuhmann D;Ananiadou S
通讯作者: Ananiadou S
DOI: 10.1136/jamia.2009.001560
发表时间: 2010-09-01
影响因子: 6.4
作者:
Savova, Guergana K.;Masanz, James J.;Chute, Christopher G.
通讯作者: Chute, Christopher G.
DOI: 10.1197/jamia.m1176
发表时间: 2003-07-01
影响因子: 6.4
作者:
Denny, JC;Smithers, JD;Spickard, A
通讯作者: Spickard, A
DOI: 10.1186/1471-2105-11-492
发表时间: 2010-09-29
期刊: BMC bioinformatics
影响因子: 3
作者:
Cohen KB;Johnson HL;Verspoor K;Roeder C;Hunter LE
通讯作者: Hunter LE