Tight data-robust bounds to mutual information combining shuffling and model selection techniques

Tight data-robust bounds to mutual information combining shuffling and model selection techniques
复制标题

DOI:
10.1162/neco.2007.19.11.2913
复制
发表时间:
2007-11-01
期刊:
影响因子:
2.9
通讯作者:
Panzeri, S.
Panzeri, S.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Montemurro, M. A.;Senatore, R.;Panzeri, S.

文献摘要

被引文献

相似文献

对脉冲时间所携带的信息的估计对于定量理解大脑功能是至关重要的,但由于有限的实验采样导致的向上偏差,这是困难的。我们提出了新的进展,基于两个基本见解,在减少偏差问题。首先,我们表明,通过仔细应用数据洗牌技术,可以几乎完全消除噪声熵的偏差,这是信息中偏差最大的部分。该方法提供了一种新的信息估计器,它比标准的直接估计器偏差小得多,并且具有相似的方差。其次,我们使用非参数检验来确定是否所有由尖峰序列编码的信息都可以在低维响应模型下被解码。如果是这种情况,响应空间的复杂性可以通过少量容易采样的参数完全捕获。结合这两种不同的方法,我们得到了一类新的信息量的精确估计,它可以提供互信息的数据鲁棒上下界。即使每个刺激的试验次数比可能的反应次数少一个数量级,这些界限也很紧。通过模拟数据和体感觉皮层的记录,测试了这些方法的有效性和实用性。该应用表明,即使存在强相关性,我们的方法也精确地限制了体内记录的真实尖峰序列编码的信息量。
The estimation of the information carried by spike times is crucial for a quantitative understanding of brain function, but it is difficult because of an upward bias due to limited experimental sampling. We present new progress, based on two basic insights, on reducing the bias problem. First, we show that by means of a careful application of data-shuffling techniques, it is possible to cancel almost entirely the bias of the noise entropy, the most biased part of information. This procedure provides a new information estimator that is much less biased than the standard direct one and has similar variance. Second, we use a nonparametric test to determine whether all the information encoded by the spike train can be decoded assuming a low-dimensional response model. If this is the case, the complexity of response space can be fully captured by a small number of easily sampled parameters. Combining these two different procedures, we obtain a new class of precise estimators of information quantities, which can provide data-robust upper and lower bounds to the mutual information. These bounds are tight even when the number of trials per stimulus available is one order of magnitude smaller than the number of possible responses. The effectiveness and the usefulness of the methods are tested through applications to simulated data and recordings from somatosensory cortex. This application shows that even in the presence of strong correlations, our methods constrain precisely the amount of information encoded by real spike trains recorded in vivo.