Automatic disabbreviation by using context information

Automatic disabbreviation by using context information
复制标题

使用上下文信息自动取消缩写

DOI:
--
复制
发表时间:
2001
期刊:
--
影响因子:
--
通讯作者:
T. Tokunaga
T. Tokunaga
中科院分区:
--
文献类型:
--
作者:
Akira Terada;T. Tokunaga

文献摘要

被引文献

相似文献

专有名词、缩略语和缩略语等未知词是文本处理的主要障碍。特别是,缩略语经常用于特定的领域。在本文中,我们提出了一种利用上下文信息的自动缩写方法。在过去的研究中,词典通常用于为缩写搜索缩写扩展候选。我们使用相同域名的缩写较少的文本,而不是词典。我们基于目标缩写的上下文与其扩展候选的上下文之间的相似性来计算扩展候选的似然。使用向量空间模型计算相似度,其中每个向量元素由周围的单词组成。利用航空领域的约1万篇文档进行的实验表明,该方法比以往的方法提高了10%的准确率。
Unknown words such as proper nouns, abbreviations, and acronyms are a major obstacle in text processing. In particular, abbreviations are often used in specific domains. In this paper, we propose an automatic disabbreviation method using context information. In past research, a dictionary has conventionally been used to search abbreviation expansion candidates for an abbreviation. We use an abbreviation-poor text of the same domain instead of a dictionary. We calculate the plausibility of expansion candidates based on the similarity between the context of a target abbreviation and that of its expansion candidates. The similarity is calculated using the vector space model, in which each vector element consists of surrounding words. Experiments using about 10,000 documents in the aviation domain showed that the proposed method is superior to past methods by 10% in precision.