A deep database of medical abbreviations and acronyms for natural language processing.

A deep database of medical abbreviations and acronyms for natural language processing.
复制标题

自然语言处理的医学缩写和首字母缩写的深度数据库。

DOI:
10.1038/s41597-021-00929-4
复制
发表时间:
2021-06-02
期刊:
影响因子:
9.8
通讯作者:
Vawdrey DK
Vawdrey DK
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Grossman Liu L;Grossman RH;Mitchell EG;Weng C;Natarajan K;Hripcsak G;Vawdrey DK

文献摘要

参考文献

相似文献

医学缩写词和首字母缩略词的识别、消歧和扩展对于防止自然语言处理中出现医学危险的误解至关重要。为了支持识别、消歧和扩展,我们推出了医学缩写和首字母缩略词元库存,这是一个医学缩写的深度数据库。对多个医疗保健专业和环境的八个来源清单进行系统协调,确定了 104,057 个缩写和 170,426 个相应的含义。使用最先进的机器学习对同义记录进行自动交叉映射,减少了冗余,从而简化了未来的应用。其他功能包括半自动质量控制以消除错误。元库存展示了新临床文本中缩写和含义的高度完整性或覆盖率,比第二大存储库有实质性改进(缩写覆盖率增加 6-14%;含义覆盖率增加 28-52%)。据我们所知,元清单是迄今为止最完整的美式英语医学缩写和首字母缩略词汇编。多种来源和高覆盖率支持在不同专业和环境中的应用。这允许跨机构的自然语言处理,这是以前的清单不支持的。元清单可在 https://bit.ly/github-clinical-abbreviations 上获取。描述报告数据的机器可访问元数据文件:10.6084/m9.figshare.14068949
The recognition, disambiguation, and expansion of medical abbreviations and acronyms is of upmost importance to prevent medically-dangerous misinterpretation in natural language processing. To support recognition, disambiguation, and expansion, we present the Medical Abbreviation and Acronym Meta-Inventory, a deep database of medical abbreviations. A systematic harmonization of eight source inventories across multiple healthcare specialties and settings identified 104,057 abbreviations with 170,426 corresponding senses. Automated cross-mapping of synonymous records using state-of-the-art machine learning reduced redundancy, which simplifies future application. Additional features include semi-automated quality control to remove errors. The Meta-Inventory demonstrated high completeness or coverage of abbreviations and senses in new clinical text, a substantial improvement over the next largest repository (6–14% increase in abbreviation coverage; 28–52% increase in sense coverage). To our knowledge, the Meta-Inventory is the most complete compilation of medical abbreviations and acronyms in American English to-date. The multiple sources and high coverage support application in varied specialties and settings. This allows for cross-institutional natural language processing, which previous inventories did not support. The Meta-Inventory is available at https://bit.ly/github-clinical-abbreviations. Machine-accessible metadata file describing the reported data: 10.6084/m9.figshare.14068949
DOI: 10.1136/postgradmedj-2016-134086
发表时间: 2016-12-01
影响因子: 5.1
作者:
Awan, Safia;Abid, Shahab;Hamid, Saeed
通讯作者: Hamid, Saeed
DOI: 10.1186/1471-2105-12-223
发表时间: 2011-06-02
期刊: BMC bioinformatics
影响因子: 3
作者:
Jimeno-Yepes AJ;McInnes BT;Aronson AR
通讯作者: Aronson AR
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.5694/mja15.00224
发表时间: 2015-08-03
影响因子: 11.4
作者:
Chemali, Mark;Hibbert, Emily J.;Sheen, Adrian
通讯作者: Sheen, Adrian
DOI: 10.1093/jamia/ocaa056
发表时间: 2020-05-29
期刊: Journal of the American Medical Informatics Association : JAMIA
影响因子: --
作者:
Lu CJ;Payne A;Mork JG
通讯作者: Mork JG