Semantic and Generative Models for Lossy Text Compression

Semantic and Generative Models for Lossy Text Compression
复制标题

有损文本压缩的语义和生成模型

DOI:
10.1093/comjnl/37.2.83
复制
发表时间:
1994
期刊:
Comput. J.
影响因子:
--
通讯作者:
H. Thimbleby
H. Thimbleby
中科院分区:
--
文献类型:
--
作者:
I. Witten;T. Bell;Alistair Moffat;C. Nevill;T. Smith;H. Thimbleby

文献摘要

被引文献

相似文献

文本压缩和图像压缩的互补范式表明,可能存在针对另一个域开发的方法的潜力。在图像编码中,有损技术产生的压缩因子比最佳无损方案的压缩因素要优越得多,我们表明文本也是如此。本文研究了传输的主观质量与其压缩因子之间的折衷。描述了两种不同的方法,可以将它们合并为一种非常有效的技术,该技术提供了比目前的最新状态更好的压缩,但可以保留原始文本和接收的文本之间合理的感知匹配程度。有损文本压缩的主要挑战是对这场比赛质量的定量评估。
The complementary paradigms of text compression and image compression suggest that there may be potential for applying methods developed for one domain to the other. In image coding, lossy techniques yield compression factors that are vastly superior to those of the best lossless schemes, and we show that this is also the case for text. This paper investigates the resulting tradeoff between subjective quality of the transmission and its compression factor. Two different methods are described, which can be combined into an extremely effective technique that provides far better compression than the present state of the art and yet preserves a reasonable degree of perceived match between the original and received text. The major challenge for lossy text compression is the quantitative evaluation of the quality of this match.