An Approach for XML Similarity Join Using Tree Serialization

An Approach for XML Similarity Join Using Tree Serialization
复制标题

DOI:
10.1007/978-3-540-78568-2_47
复制
发表时间:
2008-03
期刊:
--
影响因子:
--
通讯作者:
Lianzi Wen;Toshiyuki Amagasa;H. Kitagawa
Lianzi Wen;Toshiyuki Amagasa;H. Kitagawa
中科院分区:
其他
文献类型:
--
作者:
Lianzi Wen;Toshiyuki Amagasa;H. Kitagawa

文献摘要

相似文献

提出了一种基于XML数据序列化和后续的XML节点子序列相似性匹配的XML数据相似性连接方案。随着最近XML的爆炸性传播,现在大量的电子数据都用XML进行了标记。因此,越来越多的XML数据表示相似的内容,但具有不同的结构。为了从这些异质信息中提取尽可能多的信息,使用了相似连接。我们提出的针对XML数据的相似连接可以概括为:1)将XML数据序列化为XML节点序列;2)提取语义/结构一致的子序列;3)利用文本信息过滤出相异子序列;4)通过结构相似性检查提取子序列对作为最终结果。上述过程的执行成本很高。为了使其对大型文档集具有可伸缩性,我们使用Bloom Filter来加速文本相似度计算。通过实验验证了该方案的可行性。
This paper proposes a scheme for similarity join over XML data based on XML data serialization and subsequent similarity matching over XML node subsequences. With the recent explosive diffusion of XML, great volumes of electronic data are now marked up with XML. As a consequence, a growing amount of XML data represents similar contents, but with dissimilar structures. To extract as much information as possible from this heterogeneous information, similarity join has been used. Our proposed similarity join for XML data can be summarized as follows: 1) we serialize XML data as XML node sequences; 2) we extract semantically/structurally coherent subsequences; 3) we filter out dissimilar subsequences using textual information; and 4) we extract pairs of subsequences as the final result by checking structural similarity. The above process is costly to execute. To make it scalable against large document sets, we use Bloom filter to speed up text similarity computation. We show the feasibility of the proposed scheme by experiments.