Automatic identification and back-transliteration of foreign words for information retrieval

Automatic identification and back-transliteration of foreign words for information retrieval
复制标题

信息检索外来词自动识别与反译

DOI:
--
复制
发表时间:
1999
影响因子:
8.6
通讯作者:
K. Choi
K. Choi
中科院分区:
计算机科学1区
文献类型:
--
作者:
K. Jeong;Sung;J. Lee;K. Choi

文献摘要

被引文献

相似文献

韩语文本中出现了大量的外来语和英语词汇,特别是在科学和工程领域。我们认识到两个问题有关的外来词,这应该解决的信息检索(IR)。首先,由于外来词是动态引入的,并且不总是在字典中找到,因此它们在索引所需的形态分析中引起问题。第二,虽然一个外来词和它在源语言(如英语)中的起源指的是同一个概念,但它们被错误地视为独立的索引项。作为缓解第一个问题的一种方法,我们开发了一种算法,首先识别包含外国词的短语,然后根据统计信息从短语中提取外国词部分。对于第二个问题,我们提出了我们的方法回音译一个外来词到它的英语起源。最后,我们报告我们的评估结果为每个算法和实验结果,其对IR有效性的影响。
Many foreign words and English words appear in Korean texts, especially in the areas of science and engineering. We recognize two issues related to foreign words, which should be addressed for information retrieval (IR). First, since foreign words are introduced dynamically and not always found in a dictionary, they cause problems in morphological analysis required for indexing. Second, although a foreign word and its origin in the source language like English refer to the same concept, they are erroneously treated as independent index terms. As a way of alleviating the first problem we developed an algorithm that first identifies a phrase containing a foreign word and then extracts the foreign word part from the phrase based on statistical information. For the second problem, we present our method for back-transliteration of a foreign word to its English origin. Finally we report our evaluation results for each of the algorithms and experimental results for their impact on IR effectiveness.