Creating an online dictionary of abbreviations from MEDLINE

Creating an online dictionary of abbreviations from MEDLINE
复制标题

DOI:
10.1197/jamia.m1139
复制
发表时间:
2002-11-01
影响因子:
6.4
通讯作者:
Altman, RB
Altman, RB
中科院分区:
管理学2区
文献类型:
--
作者:
Chang, JT;Sch端tze, H;Altman, RB

文献摘要

被引文献

相似文献

目标。生物医学文献的增长对人类读者和自动算法都提出了特殊的挑战。其中一个挑战来自于文献中对缩略语的普遍和不加控制的使用。每增加一个缩写都会增加一个字段的词汇表的有效大小。因此,为了创建自动生成和维护的缩略语词典,我们开发了一种算法来匹配文本中的缩略语及其扩展。设计。我们的方法使用统计学习算法Logistic回归,根据缩略语扩展与人类注释的缩略语训练集的相似性来对缩略语扩展进行评分。我们将其应用于Medstract,这是一个MEDLINE摘要语料库,其中缩写及其扩展都已被手动标注。然后,我们在MEDLINE的所有摘要上运行该算法,创建了一本生物医学缩略语词典。为了测试数据库的覆盖率,我们使用了一个独立创建的中国医学法庭缩略语列表。我们测量了该算法在从Medstract语料库中识别缩略语的召回率和精确度。我们还测量了在数据库中检索《中国医学论坛》中的缩略语时的召回率。在Medstract语料库上,我们的算法在80%的准确率下达到了83%的召回率。将该算法应用于所有MEDLINE,得到了一个包含781,632个高分缩写的数据库。在《中国医学论坛》收录的所有缩略语中,88%在数据库中。我们开发了一种从文本中识别缩略语的算法。我们将其作为公共缩写服务器提供给\url{http://abbreviation.stanford.edu/}。
Objective. The growth of the biomedical literature presents special challenges for both human readers and automatic algorithms. One such challenge derives from the common and uncontrolled use of abbreviations in the literature. Each additional abbreviation increases the effective size of the vocabulary for a field. Therefore, to create an automatically generated and maintained lexicon of abbreviations, we have developed an algorithm to match abbreviations in text with their expansions.Design. Our method uses a statistical learning algorithm, logistic regression, to score abbreviation expansions based on their resemblance to a training set of human-annotated abbreviations. We applied it to Medstract, a corpus of MEDLINE abstracts in which abbreviations and their expansions have been manually annotated. We then ran the algorithm on all abstracts in MEDLINE, creating a dictionary of biomedical abbreviations. To test the coverage of the database, we used an independently created list of abbreviations from the China Medical Tribune.Measurements. We measured the recall and precision of the algorithm in identifying abbreviations from the Medstract corpus. We also measured the recall when searching for abbreviations from the China Medical Tribune against the database.Results. On the Medstract corpus, our algorithm achieves up to 83% recall at 80% precision. Applying the algorithm to all of MEDLINE yielded a database of 781,632 high-scoring abbreviations. Of all the abbreviations in the list from the China Medical Tribune, 88% were in the database.Conclusion. We have developed an algorithm to identify abbreviations from text. We are making this available as a public abbreviation server at \url{http: / /abbreviation.stanford.edu/}.