Building an abbreviation dictionary using a term recognition approach

Building an abbreviation dictionary using a term recognition approach
复制标题

DOI:
10.1093/bioinformatics/btl534
复制
发表时间:
2006-12-15
期刊:
影响因子:
5.8
通讯作者:
Ananiadou, Sophia
Ananiadou, Sophia
中科院分区:
生物学3区
文献类型:
--
作者:
Okazaki, Naoaki;Ananiadou, Sophia

文献摘要

被引文献

相似文献

动机:缩略语是一种高效的术语变体,需要一个缩略语词典来建立缩略语与其扩展形式之间的关联。结果:我们提出了一种识别文本集合中的缩略语定义的新方法。假设频繁出现的带有插入语的单词序列是潜在的扩展形式,我们的方法以类似于统计术语识别任务的方式识别首字母缩略词定义。该系统应用于整个MEDLINE(7811582摘要),在合理的时间内抽取了886755个候选缩略语,识别出300954个扩展形式。我们的方法优于基线系统,在我们大致模拟整个MEDLINE的评估语料库上实现了99%的准确率和82-95%的召回率。可用性和补充信息:实施和补充信息可在我们的网站上获得:http://www.chokkan.org/research/acromine/Contact:Okazaki@mi.ci.i.u-tokyo.ac.jp
Motivation: Acronyms result from a highly productive type of term variation and trigger the need for an acronym dictionary to establish associations between acronyms and their expanded forms.Results: We propose a novel method for recognizing acronym definitions in a text collection. Assuming a word sequence co-occurring frequently with a parenthetical expression to be a potential expanded form, our method identifies acronym definitions in a similar manner to the statistical term recognition task. Applied to the whole MEDLINE (7811 582 abstracts), the implemented system extracted 886 755 acronym candidates and recognized 300 954 expanded forms in reasonable time. Our method outperformed base-line systems, achieving 99% precision and 82-95% recall on our evaluation corpus that roughly emulates the whole MEDLINE.Availability and Supplementary information: The implementations and supplementary information are available at our web site: http://www.chokkan.org/research/acromine/Contact: okazaki@mi.ci.i.u-tokyo.ac.jp