Supporting efficient query processing on compressed XML files

Supporting efficient query processing on compressed XML files
复制标题

DOI:
10.1145/1066677.1066827
复制
发表时间:
2005-03
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
Yongjing Lin;Youtao Zhang;Quanzhong Li;Jun Yang
Yongjing Lin;Youtao Zhang;Quanzhong Li;Jun Yang
中科院分区:
其他
文献类型:
--
作者:
Yongjing Lin;Youtao Zhang;Quanzhong Li;Jun Yang

文献摘要

被引文献

相似文献

XML已经被广泛接受为数据表示和交换的事实上的格式。然而,它也被称为在其表示的信息冗余过多。虽然已经提出了各种压缩方案,其中一些可以支持查询处理压缩文件,它通常是不可避免的执行部分(或全部)数据解压缩,这是昂贵的,在某些情况下,可能会占主导地位的查询处理时间。通过将压缩结果组织成一组上下文无关的语法规则,该方案支持在不解压缩的情况下高效地处理XML查询。实验结果表明,该方案在查询处理时间上优于gzip算法,压缩比与gzip相当。
XML has been widely accepted as the de facto format for data representation and exchange. However, it is also known for the excessive information redundancy in its representation. While various compression schemes have been proposed and some of them can support query processing over compressed files, it is usually inevitable to perform partial (or full) data decompression which is expensive and in some cases may dominate the query processing time.In this paper, we propose a new XML compression scheme based on the Sequitur compression algorithm. By organizing the compression result as a set of context free grammar rules, the scheme supports efficient processing of XPath queries without decompression. The experimental results show that this scheme achieves comparable compression ratio as gzip while its query processing time is among the best of existing algorithms.