On Optimal Data Compression in Multiterminal Statistical Inference

On Optimal Data Compression in Multiterminal Statistical Inference
复制标题

多端统计推断中的最优数据压缩

DOI:
10.1109/tit.2011.2162270
复制
发表时间:
2011
影响因子:
2.5
通讯作者:
S. Amari
S. Amari
中科院分区:
计算机科学2区
文献类型:
--
作者:
S. Amari

文献摘要

被引文献

相似文献

统计推理的多端理论涉及估计或测试在每个源的一定传输速率的限制下,从两个(或多个)相关信息源产生的字母相关性的问题。一个典型的示例是两个具有关节概率<i> p </i>(<i> x </i>,<i> y </i>)的二进制源,其中<i> x </i>的相关性和<i> y </i>将进行测试或估计。给定<i> n </i> iid观测<i> x </i> <sup> n </sup> = <i> x </i> <sub> 1 </sub> ... <i> x </i> <sub> n </sub>和<i> y </i> <sup> n </sup> = <i> y </i> <sub> 1 </sub> ... <i> y </i> <sup> n </sup>,唯一<; 1)每个位都可以传输到共同的目的地。统计推断的最佳数据压缩是什么?一个简单的想法是发送第一个<i> k <i> x </i> <sup> n </sup>和<i> y </i> <sup> n </i> >。一个更简单的问题是助手情况,在所有<i> y> y </i> <sup> n </sup>是传输的。确定是否有更好的数据压缩方案是一个长期存在的问题,而不是发送第一个<i> k </i>字母的简单方案。本文在线性 - 阈值编码的框架下搜索最佳数据压缩,并表明根据相关值有更好的数据压缩方案。为此,我们评估了线性阈值压缩方案类别中的Fisher信息。还可以证明,当x </i>和<i> y </i>是独立的,或者它们的相关性不太大时,简单方案是最佳的。
The multiterminal theory of statistical inference deals with the problem of estimating or testing the correlation of letters generated from two (or many) correlated information sources under the restriction of a certain transmission rate for each source. A typical example is two binary sources with joint probability <i>p</i>(<i>x</i>, <i>y</i>) where the correlation of <i>x</i> and <i>y</i> is to be tested or estimated. Given <i>n</i> iid observations <i>x</i><sup>n</sup> = <i>x</i><sub>1</sub> ...<i>x</i><sub>n</sub> and <i>y</i><sup>n</sup>=<i>y</i><sub>1</sub> ...<i>y</i><sup>n</sup>, only <i>k</i> = <i>rn</i> (0 <; <i>r</i> <; 1) bits each can be transmitted to a common destination. What is the optimal data compression for statistical inference? A simple idea is to send the first <i>k</i> letters of <i>x</i><sup>n</sup> and <i>y</i><sup>n</sup>. A simpler problem is the helper case where the optimal data compression of <i>x</i><sup>n</sup> is searched for under the condition that all of <i>y</i><sup>n</sup> are transmitted. It is a long standing problem to determine if there is a better data compression scheme than this simple scheme of sending first <i>k</i> letters. The present paper searches for the optimal data compression under the framework of linear-threshold encoding and shows that there is a better data compression scheme depending on the value of correlation. To this end, we evaluate the Fisher information in the class of linear-threshold compression schemes. It is also proved that the simple scheme is optimal when <i>x</i> and <i>y</i> are independent or their correlation is not too large.