FINDING APPROXIMATE MATCHES IN LARGE LEXICONS

FINDING APPROXIMATE MATCHES IN LARGE LEXICONS
复制标题

DOI:
10.1002/spe.4380250307
复制
发表时间:
1995-03-01
影响因子:
3.5
通讯作者:
DART, P
DART, P
中科院分区:
计算机科学4区
文献类型:
--
作者:
ZOBEL, J;DART, P

文献摘要

被引文献

相似文献

近似字符串匹配用于拼写校正和人名匹配。在本文中,我们将展示如何使用字符串匹配技术结合词典索引,找到近似匹配的大型词典。我们测试了几种词典索引技术,包括n元语法和排列词典,以及几种字符串匹配技术,包括字符串相似性度量和语音编码。我们提出了这些技术相结合的方法,并通过实验表明,这些组合产生良好的检索效果,同时保持索引大小和检索时间低。我们的实验还表明,在以前的索赔相比,语音编码是显着劣于字符串距离的措施,这被证明是适合于拼写校正和个人姓名匹配。
Approximate string matching is used for spelling correction and personal name matching. In this paper we show how to use string matching techniques in conjunction with lexicon indexes to find approximate matches in a large lexicon. We test several lexicon indexing techniques, including n-grams and permuted lexicons, and several string matching techniques, including string similarity measures and phonetic coding. We propose methods for combining these techniques, and show experimentally that these combinations yield good retrieval effectiveness while keeping index size and retrieval time low. Our experiments also suggest that, in contrast to previous claims, phonetic codings are markedly inferior to string distance measures, which are demonstrated to be suitable for both spelling correction and personal name matching.