Optimal structure of metaplasticity for adaptive learning.

Optimal structure of metaplasticity for adaptive learning.
复制标题

自适应学习的最佳结构。

DOI:
10.1371/journal.pcbi.1005630
复制
发表时间:
2017-06
影响因子:
4.3
通讯作者:
Soltani A
Soltani A
中科院分区:
生物学2区
文献类型:
--
作者:
Khorsand P;Soltani A

文献摘要

参考文献

被引文献

相似文献

在不断变化的环境中学习奖励反馈需要高度的适应性,但对奖励信息的准确估计需要缓慢的更新。在估计奖赏概率的框架内,我们研究了如何通过元可塑性来缓解这种适应性和精确性之间的权衡,即突触的变化并不总是改变突触的有效性。使用平均场和蒙特卡罗模拟,我们确定了可以实质上克服适应性和精度之间的权衡的“优越的”化塑性模型。这些模型可以通过形成两组独立的元状态来实现适应性和精确度:储备库和缓冲器。处于储备性元状态的突触不会因奖赏反馈而改变其有效性,而处于缓冲元状态的突触则可以改变其有效性。功效的快速变化仅限于突触占用缓冲区,造成了一个瓶颈,既能减少噪音,又不会显著降低适应性。相比之下,人口较多的水库可以产生强烈的信号,而不表现出任何可观察到的可塑性。通过比较我们的模型和几个竞争模型在动态概率估计任务中的行为,我们发现在更大范围的模型参数范围内,优越的化塑性模型的性能接近最优。最后,我们发现元塑料模型对模型参数的变化是健壮的,而且元塑料过渡对于自适应学习是至关重要的,因为用分级塑料过渡(改变突触有效性的过渡)取代它们会降低克服适应性和精度权衡的能力。总体而言,我们的结果表明,突触变化的无处不在的不可靠性表现出可塑性,这可以提供一种稳健的机制来缓解适应性和精确度之间的权衡,从而减轻适应性学习。成功地从我们的经验和环境反馈中学习,需要在每次反馈后以精确的数量更新分配给给定选项或行动的奖励值。在被称为强化学习的基于奖励的学习的标准模型中,学习率决定了这种更新的强度。较大的学习率允许快速更新值(适应性较强),但会引入噪声(精度较小),而较小的学习率则相反。因此,学习似乎受到适应性和精确度之间的权衡。在这里,我们询问是否有突触机制能够根据奖励统计数据调整大脑的可塑性水平,从而允许学习过程具有适应性。我们表明,突触状态的变化,即突触状态的变化,塑造了未来的突触修改,而突触的强度没有任何明显的变化,我们证明了这种变化可以提供这样的机制,并进一步确定了这种元可塑性的最佳结构。我们认为,元可塑性有时不会导致可观察到的行为变化,因此可以被认为是缺乏学习,它可以为适应性学习提供一种稳健的机制。
Learning from reward feedback in a changing environment requires a high degree of adaptability, yet the precise estimation of reward information demands slow updates. In the framework of estimating reward probability, here we investigated how this tradeoff between adaptability and precision can be mitigated via metaplasticity, i.e. synaptic changes that do not always alter synaptic efficacy. Using the mean-field and Monte Carlo simulations we identified ‘superior’ metaplastic models that can substantially overcome the adaptability-precision tradeoff. These models can achieve both adaptability and precision by forming two separate sets of meta-states: reservoirs and buffers. Synapses in reservoir meta-states do not change their efficacy upon reward feedback, whereas those in buffer meta-states can change their efficacy. Rapid changes in efficacy are limited to synapses occupying buffers, creating a bottleneck that reduces noise without significantly decreasing adaptability. In contrast, more-populated reservoirs can generate a strong signal without manifesting any observable plasticity. By comparing the behavior of our model and a few competing models during a dynamic probability estimation task, we found that superior metaplastic models perform close to optimally for a wider range of model parameters. Finally, we found that metaplastic models are robust to changes in model parameters and that metaplastic transitions are crucial for adaptive learning since replacing them with graded plastic transitions (transitions that change synaptic efficacy) reduces the ability to overcome the adaptability-precision tradeoff. Overall, our results suggest that ubiquitous unreliability of synaptic changes evinces metaplasticity that can provide a robust mechanism for mitigating the tradeoff between adaptability and precision and thus adaptive learning. Successful learning from our experience and feedback from the environment requires that the reward value assigned to a given option or action to be updated by a precise amount after each feedback. In the standard model for reward-based learning known as reinforcement learning, the learning rates determine the strength of such update. A large learning rate allows fast update of values (large adaptability) but introduces noise (small precision), whereas a small learning rate does the opposite. Thus, learning seems to be bounded by a tradeoff between adaptability and precision. Here, we asked whether there are synaptic mechanisms that are capable of adjusting the brain’s level of plasticity according to reward statistics, and, therefore, allow the learning process to be adaptive. We showed that metaplasticity, changes in the synaptic state that shape future synaptic modifications without any observable changes in the strength of synapses, could provide such a mechanism and furthermore, identified the optimal structure of such metaplasticity. We propose that metaplasticity, which sometimes causes no observable changes in behavior and thus could be perceived as a lack of learning, can provide a robust mechanism for adaptive learning.
DOI: 10.1038/nn.2250
发表时间: 2009-02
影响因子: 25
作者:
Moussawi, Khaled;Pacchioni, Alejandra;Moran, Megan;Olive, M. Foster;Gass, Justin T.;Lavin, Antonieta;Kalivas, Peter W.
通讯作者: Kalivas, Peter W.
DOI: 10.1038/nn1859
发表时间: 2007-04-01
影响因子: 25
作者:
Fusi, Stefano;Abbott, L. F.
通讯作者: Abbott, L. F.
DOI: 10.1523/jneurosci.1989-14.2015
发表时间: 2015-02-11
影响因子: 5.3
作者:
Costa, Vincent D.;Tran, Valery L.;Averbeck, Bruno B.
通讯作者: Averbeck, Bruno B.
DOI: 10.1016/j.neuron.2009.02.026
发表时间: 2009-04-30
期刊: NEURON
影响因子: 16.2
作者:
Enoki, Ryosuke;Hu, Yi-ling;Fine, Alan
通讯作者: Fine, Alan
DOI: 10.1016/0024-3795(86)90210-7
发表时间: 1986-04-01
影响因子: 1.1
作者:
FUNDERLIC, RE;MEYER, CD
通讯作者: MEYER, CD