An Approach for XML Similarity Join Using Tree Serialization
An Approach for XML Similarity Join Using Tree Serialization
复制标题
DOI:
10.1007/978-3-540-78568-2_47
复制
发表时间:
2008-03
期刊:
影响因子:
--
通讯作者:
Lianzi Wen;Toshiyuki Amagasa;H. Kitagawa
中科院分区:
文献类型:
--
作者:
Lianzi Wen;Toshiyuki Amagasa;H. Kitagawa
This paper proposes a scheme for similarity join over XML data based on XML data serialization and subsequent similarity matching over XML node subsequences. With the recent explosive diffusion of XML, great volumes of electronic data are now marked up with XML. As a consequence, a growing amount of XML data represents similar contents, but with dissimilar structures. To extract as much information as possible from this heterogeneous information, similarity join has been used. Our proposed similarity join for XML data can be summarized as follows: 1) we serialize XML data as XML node sequences; 2) we extract semantically/structurally coherent subsequences; 3) we filter out dissimilar subsequences using textual information; and 4) we extract pairs of subsequences as the final result by checking structural similarity. The above process is costly to execute. To make it scalable against large document sets, we use Bloom filter to speed up text similarity computation. We show the feasibility of the proposed scheme by experiments.