Self-supervised autoencoder framework for salient sensory feature extraction
Self-supervised autoencoder framework for salient sensory feature extraction
批准号:
2894189
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
自然界充满了噪音,但大脑传递信息的能力却受到严重限制。因此,丢弃感官输入中包含的无关信息,同时保留与输入标签相关的显著特征,是生存的关键。许多研究表明,这可以部分通过大脑实现信息瓶颈来实现,其目的是最大化压缩和保存重要信息之间的权衡。实现信息瓶颈需要一种测量信息的方法。神经科学中常用的度量是信息理论度量-互信息,它描述了神经元反应可以告诉我们的关于刺激的信息量。然而,该度量和许多现有的度量估计方法都受到维度诅咒的影响。近年来,这一挑战已经通过将互信息的估计框架化为对抗设置中的最小最大优化问题来解决,这需要两个神经网络玩一个游戏,其中一个网络性能的提高导致另一个网络性能的恶化。与传统的互信息估计方法不同,该方法在维数和样本量上具有可扩展性。在这项研究的基础上,我们提出了一个对抗启发的显著感官特征提取自编码器框架。该框架由三个神经网络组成:编码器、解码器和分类器。编码器的目标是学习压缩数据,使分类器能够准确地对输入进行分类,但解码器无法完全重建原始输入。辅助分类任务的实现是为了帮助调节编码器的潜在空间以捕获显著特征。通过在MNIST数据集和CIFAR10数据集的子集上对框架进行初步训练,结果表明该框架将图像数据中的不相关信息丢弃。此外,它似乎执行了图像-背景分离,这种现象通过将视觉场景分割成物体和背景区域来感知形状和物体。在这个项目中,我们的目标是通过在更复杂的图像数据集上训练框架来证实这些发现,并研究它是否可以重现视觉信息处理中实现的其他特征选择机制。我们还旨在将该框架的应用扩展到其他感官模式,即语音处理。人类听众有一种非凡的能力,可以从嘈杂的背景中分辨出说话者的讲话——这种现象被称为鸡尾酒会问题。然而,我们对哪些信息被丢弃的理解是有限的。生理学研究经常使用来自其他领域的特征,这使得它们很难与听觉神经生理学联系起来。同样,建模研究通常依赖于注释的语言特征空间,而不是利用语音的光谱时间动态。后者已被证明可以改善模型中的特征提取。因此,通过使用直接从输入数据中获得的表征,我们的框架可以为神经对刺激的反应提供更精确的解释。第一步将是训练框架完成一项与语音处理相关的简单任务,比如对不同的元音进行分类,并将学习到的特征与听觉神经生理学的现有文献进行比较。然后,我们的研究结果可以用来产生新的、可测试的假设,这些假设是关于大脑中抗噪声语音处理的基本特征。此外,该框架所学习的特征也可用于改进语音识别系统。
英文摘要
The natural world is full of noise, but the brain's capacity for information transmission is severely limited. Therefore, discarding irrelevant information contained in sensory inputs while retaining salient features, which are related to the input label, is key to survival. Numerous studies suggest that this could be partly achieved by the brain implementing information bottlenecks, which aim to maximise the trade-off between compression and preservation of salient information. Implementation of information bottlenecks requires a way of measuring information. A commonly adopted measure in neuroscience is the information theoretic metric - mutual information, which describes the amount of information a neuronal response can tell us about a stimulus. However, this metric and many existing approaches for its estimation suffer from the curse of dimensionality.In recent years, this challenge has been approached by framing estimation of mutual information as a minmax optimisation problem in an adversarial setting, which entails two neural networks playing a game in which the improvement in one network's performance leads to the worsening of the other's. This method has been found to be scalable in dimension and sample-size unlike traditional mutual information estimators. Building on this research, we propose an adversarial-inspired autoencoder framework for salient sensory feature extraction. The proposed framework consists of three neural networks: the encoder, the decoder, and the classifier. The objective is for the encoder to learn to compress the data such that the classifier can accurately classify the input, but the decoder cannot fully reconstruct the original input. The auxiliary classification task is implemented to help condition the latent space of the encoder to capture salient features. Preliminary results obtained by training the framework on the MNIST dataset and subsets of the CIFAR10 dataset show that the framework discards irrelevant information from image data. Furthermore, it appears that it performs figure-ground separation, a phenomenon that enables perception of shapes and objects by segmenting visual scenes into object- and background-like regions. In this project, we aim to confirm these findings by training the framework on more complex image datasets and investigate whether it can reproduce other feature selective mechanisms implemented in visual information processing. We also aim to extend the application of the framework to other sensory modalities, namely to speech processing. Human listeners have the remarkable ability to separate the speech of one speaker from a noisy background - a phenomenon called the cocktail party problem. However, our understanding of what information is discarded is limited. Physiological studies often make use of features derived from other fields, making them difficult to relate to auditory neurophysiology. Similarly, modelling studies often rely on annotated linguistic feature spaces rather than utilising the spectrotemporal dynamics of speech. The latter has been demonstrated to improve feature extraction in models. Therefore, by using representations obtained directly from input data, our framework may provide a more precise explanation of neural responses to stimuli. The first step will be to train the framework on a simple task related to speech processing, such as classifying different vowels, and comparing the features learned to existing literature on auditory neurophysiology. Our results could then be used to generate new, testable hypotheses about the features underlying noise-robust speech processing in the brain. Furthermore, features learned by the framework could also be used to improve speech recognition systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于指点触控行为的身份认证与监控方法研究
-
批准号:61175039
-
项目类别:面上项目
-
资助金额:59.0万元
-
批准年份:2011
-
负责人:蔡忠闽
-
依托单位: