FINDING APPROXIMATE MATCHES IN LARGE LEXICONS
FINDING APPROXIMATE MATCHES IN LARGE LEXICONS
复制标题
DOI:
10.1002/spe.4380250307
复制
发表时间:
1995-03-01
影响因子:
3.5
通讯作者:
DART, P
中科院分区:
文献类型:
--
作者:
ZOBEL, J;DART, P
Approximate string matching is used for spelling correction and personal name matching. In this paper we show how to use string matching techniques in conjunction with lexicon indexes to find approximate matches in a large lexicon. We test several lexicon indexing techniques, including n-grams and permuted lexicons, and several string matching techniques, including string similarity measures and phonetic coding. We propose methods for combining these techniques, and show experimentally that these combinations yield good retrieval effectiveness while keeping index size and retrieval time low. Our experiments also suggest that, in contrast to previous claims, phonetic codings are markedly inferior to string distance measures, which are demonstrated to be suitable for both spelling correction and personal name matching.