Likelihood-Based Data Squashing: A Modeling Approach to Instance Construction

Likelihood-Based Data Squashing: A Modeling Approach to Instance Construction
复制标题

基于似然的数据压缩:实例构建的建模方法

DOI:
--
复制
发表时间:
2002
影响因子:
4.8
通讯作者:
G. Ridgeway
G. Ridgeway
中科院分区:
计算机科学3区
文献类型:
--
作者:
D. Madigan;Nandini Raghavan;W. DuMouchel;M. Nason;C. Posse;G. Ridgeway

文献摘要

被引文献

相似文献

压缩是一种保留统计信息的有损数据压缩技术。具体来说,压缩将大规模数据集压缩为小得多的数据集,使得对较小(压缩)数据集进行的统计分析的输出再现对原始数据集进行的相同统计分析的输出。基于可能性的数据压缩(LDS)与以前发布的压缩算法不同,因为它使用统计模型来压缩数据。结果表明,LDS提供了优异的挤压性能,即使目标统计分析偏离用于挤压数据的模型。
Squashing is a lossy data compression technique that preserves statistical information. Specifically, squashing compresses a massive dataset to a much smaller one so that outputs from statistical analyses carried out on the smaller (squashed) dataset reproduce outputs from the same statistical analyses carried out on the original dataset. Likelihood-based data squashing (LDS) differs from a previously published squashing algorithm insofar as it uses a statistical model to squash the data. The results show that LDS provides excellent squashing performance even when the target statistical analysis departs from the model used to squash the data.