Side effects of learning from low-dimensional data embedded in a Euclidean space

Side effects of learning from low-dimensional data embedded in a Euclidean space
复制标题

DOI:
10.1007/s40687-023-00378-y
复制
发表时间:
2022-03
影响因子:
1.2
通讯作者:
Juncai He;R. Tsai;Rachel A. Ward
Juncai He;R. Tsai;Rachel A. Ward
中科院分区:
数学3区
文献类型:
--
作者:
Juncai He;R. Tsai;Rachel A. Ward

文献摘要

被引文献

相似文献

低维流形假设假定在许多应用中发现的数据,例如涉及自然图像的数据,(近似)位于嵌入高维欧几里得空间的低维流形上。在这种情况下,典型的神经网络定义了一个函数,该函数将嵌入空间中的有限数量的向量作为输入。然而,人们通常需要考虑在训练分布之外的点处评估优化的网络。本文考虑的情况下,其中的训练数据分布在一个线性子空间。我们得到的学习函数的变化,定义的神经网络,在横向的子空间的方向上的估计。我们研究了与网络深度和数据流形余维中的噪声相关的潜在正则化效应。由于噪声的存在,我们还提出了训练中的其他副作用。
The low-dimensional manifold hypothesis posits that the data found in many applications, such as those involving natural images, lie (approximately) on low-dimensional manifolds embedded in a high-dimensional Euclidean space. In this setting, a typical neural network defines a function that takes a finite number of vectors in the embedding space as input. However, one often needs to consider evaluating the optimized network at points outside the training distribution. This paper considers the case in which the training data are distributed in a linear subspace of. We derive estimates on the variation of the learning function, defined by a neural network, in the direction transversal to the subspace. We study the potential regularization effects associated with the network’s depth and noise in the codimension of the data manifold. We also present additional side effects in training due to the presence of noise.