s-grams:: Defining generalized n-grams for information retrieval

s-grams:: Defining generalized n-grams for information retrieval
复制标题

DOI:
10.1016/j.ipm.2006.09.016
复制
发表时间:
2007-07-01
影响因子:
8.6
通讯作者:
Jarvelin, Kalervo
Jarvelin, Kalervo
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jarvelin, Anni;Jarvelin, Antti;Jarvelin, Kalervo

文献摘要

被引文献

相似文献

n-grams have been used widely and successfully for approximate string matching in many areas. s-grams have been introduced recently as an n-gram based matching technique, where di-grams are formed of both adjacent and non-adjacent characters. s-grams have proved successful in approximate string matching across language boundaries in Information Retrieval (IR). s-grams however lack precise definitions. Also their similarity comparison lacks precise definition. In this paper.. we give precise definitions for both. Our definitions are developed in a bottom-up manner, only assuming character strings and elementary mathematical concepts. Extending established practices, we provide novel definitions of s-gram profiles and the L, distance metric for them. This is a stronger string proximity measure than the popular Jaccard similarity measure because Jaccard is insensitive to the counts of each n-gram in the strings to be compared. However, due to the popularity of Jaccard in IR experiments, we define the reduction of s-gram profiles to binary profiles in order to precisely define the (extended) Jaccard similarity function for s-grams. We also show that n-gram similarity/distance computations are special cases of our generalized definitions. (c) 2006 Elsevier Ltd. All rights reserved.