Lossy Text Compression Techniques

Lossy Text Compression Techniques
复制标题

有损文本压缩技术

DOI:
10.1007/978-1-84628-992-7_28
复制
发表时间:
2007
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
S. Latifi
S. Latifi
中科院分区:
--
文献类型:
--
作者:
Venka Palaniappan;S. Latifi

文献摘要

被引文献

相似文献

大多数文本文档都包含大量的冗余。数据压缩可用于最小化这种冗余并提高传输效率或节省存储空间。已经引入了几种文本压缩算法用于关键应用领域中使用的无损文本压缩。对于非关键应用程序,我们可以使用有损文本压缩来提高压缩效率。在本文中,我们提出了三种不同的源模型,基于字符的有损文本压缩:元音脱落(DOV),字母映射(LMP)和字符替换(ROC)。介绍了这些方法的工作原理和改造方法。压缩比得到的包括和比较。并与霍夫曼编码和算术编码算法进行了性能比较。最后,提出了进一步提高性能的一些想法。
Most text documents contain a large amount of redundancy. Data compression can be used to minimize this redundancy and increase transmission efficiency or save storage space. Several text compression algorithms have been introduced for lossless text compression used in critical application areas. For non-critical applications, we could use lossy text compression to improve compression efficiency. In this paper, we propose three different source models for character-based lossy text compression: Dropped Vowels (DOV), Letter Mapping (LMP), and Replacement of Characters (ROC). The working principles and transformation methods associated with these methods are presented. Compression ratios obtained are included and compared. Comparisons of performance with those of the Huffman Coding and Arithmetic Coding algorithm are also made. Finally, some ideas for further improving the performance already obtained are proposed.