Sound Source Localization Inside a Structure Under Semi-Supervised Conditions

Sound Source Localization Inside a Structure Under Semi-Supervised Conditions
复制标题

DOI:
10.1109/taslp.2023.3263776
复制
发表时间:
2023
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
S. Kita;Y. Kajikawa
S. Kita;Y. Kajikawa
中科院分区:
其他
文献类型:
--
作者:
S. Kita;Y. Kajikawa

文献摘要

相似文献

我们提出了一种在真实环境中应用基于模拟数据训练的声源定位(SSL)模型的方法,以及针对结构内的声源定位(SSL)的域转移(DT)模型。DT模型将真实数据转换为伪仿真数据。然后,使用DT模型使在仿真数据上训练的SSL模型适应于真实数据。我们的方法由一个SSL模型和一个DT模型组成。SSL模型预测声源在结构中的位置,而DT模型则转换数据。由于我们的模拟并不完美,因此需要对实际数据进行外推,以便与SSL模型一起使用。然而,DT模型变换的数据是在特征空间内进行内插的。结果是,在现实世界中,SSL模型的性能得到了改进。在我们的研究中,在结构外表面观测到的加速度计的频谱作为模型输入。目标是预测声源的位置。使用深度和卷积神经网络构建SSL模型,使用自动编码器、深度卷积自动编码器或像素2Pix构建DT模型。T分布随机邻居嵌入的二维分布表明,使用Pix2pix作为DT模型具有最好的性能。此外,与不应用变换的情况相比,我们的方法在分类问题上的性能提高了57%,在回归问题上提高了27%。
We propose a method for applying a sound source localization (SSL) model trained on simulated data in a real-world environment, with a domain transfer (DT) model for the SSL inside a structure. The DT model transfers real data into pseudo-simulation data. The SSL model trained on the simulation data is then adapted to the real data using the DT model. Our method consists of an SSL model and a DT model. The SSL model predicts the position of a sound source inside the structure, whereas the DT model transforms the data. Because our simulation is not perfect, real data are extrapolated for use with the SSL model. However, the data transformed by the DT model are interpolated within the feature space. The outcome is that the performance of the SSL model in the real world is improved. In our study, the frequency spectra of accelerometers observed on the outer surface of the structure are the model input. The goal is to predict the position of the sound source. The SSL model is built using deep and convolutional neural networks, and the DT model is built using either an autoencoder, a deep convolutional autoencoder, or pix2pix. The two-dimensional distributions of the t-distributed Stochastic Neighbor Embedding indicate that using pix2pix as the DT model shows the best performance. Furthermore, our method's performance for SSL is improved by 57% for the classification problem and by 27% for the regression problem when compared to the case where no transformation is applied.