Auditory perception versus automatic estimation of location and orientation of an acoustic source in a real environment

Auditory perception versus automatic estimation of location and orientation of an acoustic source in a real environment
复制标题

DOI:
10.1250/ast.31.309
复制
发表时间:
2010-09
影响因子:
0.7
通讯作者:
A. Y. Nakano;S. Nakagawa;Kazumasa Yamamoto
A. Y. Nakano;S. Nakagawa;Kazumasa Yamamoto
中科院分区:
--
文献类型:
--
作者:
A. Y. Nakano;S. Nakagawa;Kazumasa Yamamoto

文献摘要

相似文献

在这项工作中,方向性声源的位置和方向的感知在一个真实的封闭的环境中由蒙眼的听众进行了调查和比较的方法,自动估计的位置和方向的源使用T形麦克风阵列。在主观实验中,使用蒙上眼睛的听众,一个人的发言者作为声源和听众判断说话者的面对角度(四个可能的方向之一,偏移90 °)和位置后,听一句话。该程序在训练阶段之前和之后进行了两次。在训练中,听众被允许去除眼罩,并验证扬声器的位置和方向。训练后,正确定向率从75.0%提高到76.5%,平均位置误差从66.2cm降低到60.6cm。此外,在同一真实的环境中,对8个方向偏移45 °的方向估计进行了主观实验,结果表明,在真实的环境中的方向估计比在无回声环境中的方向估计更困难。在自动估计方法中使用了人工神经网络。在由8个T形麦克风阵列组成的阵列网络中,最靠近被蒙住的收听者的T形麦克风阵列获得了68.1%的正确定向率和48.0cm的平均位置误差,使得能够在人类听觉感知和自动估计方法之间进行粗略比较(67%的正确定向率和38.6cm的较好平均位置误差是通过网络中的T形阵列获得的最佳结果)。有人澄清,自动估计方法不能超过听觉系统在正确的方向比,但是,它产生了更好的结果的平均位置误差。
In this work, the perception of the position and orientation of a directional acoustic source in a real enclosed environment by blindfolded listeners is investigated and compared with a method that automatically estimates the position and orientation of the source using a T-shaped microphone array. In the subjective experiment using blindfolded listeners, a human speaker acted as an acoustic source and listeners judged the speaker's facing angle (one out of four possible orientations shifted by 90 � ) and position after listening to a spoken sentence. This procedure was performed twice, before and after a training phase. In the training, listeners were allowed to remove the blindfold and verify the speaker's position and orientation. After the training, the correct orientation ratio increased from 75.0 to 76.5% and the average position error decreased from 66.2 to 60.6 cm. In addition, a subjective experiment on orientation estimation with eight orientations shifted by 45 � in the same real environment showed that orientation estimation in a real environment was more difficult than that in an anechoic environment. Artificial neural networks (ANNs) were used in the automatic estimation method. A correct orientation ratio of 68.1% and an average position error of 48.0 cm were obtained by the T-shaped microphone array located nearest to the blindfolded listener among an array network consisting of eight T-shaped microphone arrays, enabling a rough comparison between human auditory perception and the automatic estimation method (a correct orientation ratio of 67% and a better average position error of 38.6 cm were the best results obtained by a T-shaped array in the network). It was clarified that the automatic estimation method cannot surpass the auditory system in terms of correct orientation ratio; however, it yielded better results in terms of the average position error.