Browsing and searching compressed documents

Browsing and searching compressed documents
复制标题

浏览和搜索压缩文档

DOI:
--
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
R. Wan
R. Wan
中科院分区:
--
文献类型:
--
作者:
R. Wan

文献摘要

被引文献

相似文献

压缩和信息检索是文档管理的两个领域,由于实现其目标的相互冲突的方法而分开存在。本研究探讨一种机制,提供无损压缩和基于短语的浏览和搜索的大型文档集。调查的框架是一个现有的离线基于字典的压缩算法。在以前的工作和实验的支持下,对算法的分析突出了两个对检索很重要的因素:高效的解码和单独的字典流。然而,在将该算法纳入浏览系统之前,有三个方面的改进是必要的。首先,为了适应检索,算法必须产生一个建立在单词而不是字符上的字典。引入了一个预处理阶段,它将消息分为单词和非单词,沿着单词修饰符。其次,该算法的内存需求阻碍了大型文档的处理。早期的工作提出了一种解决方案,在压缩之前将消息分成单独的块。在这里,提出了一个后处理阶段,它结合了一系列阶段的块。实验表明,执行的阶段数和压缩级别的改进之间的权衡。向用户提供信息检索系统访问的组织有时关心可用磁盘空间的量。但是用户有不同的观点,可能会把响应时间放在更高的优先级。事实上,更快的计算机和网络连接转化为用户的耐心水平。压缩算法的最后一个改进是两个新的编码方案,取代了以前工作中使用的熵编码器。虽然部署它们会牺牲压缩效率,但这两种机制可以提高效率,正如实验所示。通过对压缩算法的改进,提出了一种有效支持短语浏览的技术。短语上下文可以通过单词修饰词进行搜索和逐步细化。由于对算法进行了三项更改,短语在视觉上更具吸引力,可以处理更大的文档,并且响应时间得到改善。
Compression and information retrieval are two areas of document management that exist separately due to the conflicting methods of achieving their goals. This research examines a mechanism which provides lossless compression and phrase-based browsing and searching of large document collections. The framework for the investigation is an existing off-line dictionary-based compression algorithm. An analysis of the algorithm, supported by previous work and experiments, highlights two factors that are important for retrieval: efficient decoding, and a separate dictionary stream. However, three areas of improvement are necessary, prior to the inclusion of the algorithm into a browsing system. First, in order to accommodate retrieval, the algorithm must produce a dictionary built up on words, rather than characters. A pre-processing stage is introduced which separates the message into words and non-words, along with word modifiers. Second, the memory requirements of the algorithm prevent the processing of large documents. Earlier work has proposed a solution which separates the message into individual blocks prior to compression. Here, a post-processing stage is proposed which combines the blocks in a series of phases. Experiments show the trade-offs between the number of phases performed and the improvements in compression levels. Organisations which provide access to an information retrieval system to users are sometimes concerned with the amount of disk space available. But users have a different point of view and may place response time at a higher priority. Indeed, faster computers and network connections translate into plummeting patience levels of users. The last improvement to the compression algorithm is two new coding schemes which replace the entropy coder that was used in previous work. While deploying them sacrifices compression effectiveness, these two mechanisms offer improved efficiency, as shown through experiments. With the enhancements to the compression algorithm in hand, a technique to efficiently support phrase browsing is presented. Phrase contexts can be searched and progressively refined through the word modifiers. Because of the three changes to the algorithm, phrases are more visually appealing, larger documents can be processed, and response times are improved.