Vocal Tract Length Estimation Using Accumulated Means of Formants and Its Effects on Speaker-Normalization

Vocal Tract Length Estimation Using Accumulated Means of Formants and Its Effects on Speaker-Normalization
复制标题

使用共振峰累积方法估计声带长度及其对说话人归一化的影响

DOI:
10.1109/taslp.2021.3060172
复制
发表时间:
2021
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Akira Watanabe
Akira Watanabe
中科院分区:
--
文献类型:
--
作者:
Tadashi Sakata;Naomitsu Ikeda;Yuichi Ueda;Akira Watanabe

文献摘要

相似文献

单个说话者的声道长度的差异导致音素声学特征的变化。本文提出了一种简单的方法来估计特定说话人的虚拟带库,并定量评价虚拟带库的某些说话人归一化效果。我们采用形成峰轨迹的累积方法来估计从儿童到成人的说话者的vtl。对于形成峰估计,采用了反滤波控制(IFC)系统。在系统中,分析顺序的决定,即要估计的共振峰的数量,是自动化的。此外,为了评估虚拟带库的说话人归一化效果,我们提出了数据约简方法,该方法可以合理地从构造空间的分布中找到密集的椭圆区域。使用这些椭圆区域,我们评估了虚拟带库的三种规范化效果:通过所有虚拟带库的均值作为标准进行规范化,通过虚拟带库的说话者分类方法进行规范化,以及通过单个虚拟带库进行规范化。与原始数据的标准面积相比,分类平均和单个虚拟带库的面积分别减少了39.5%和46.6%。结果,我们提出的方法被用来提供一个“标准化元音映射(NVM)”,将通用元音分布可视化为语言信息的核心图像。最后,我们将所提出的方法与基于磁共振成像(MRI)数据的另一种方法的估计带库进行了比较。
Differences in vocal tract lengths (VTLs) in individual speakers cause variations in acoustic features of phonemes. In this paper, a simple method to estimate speaker-specific VTLs and to quantitatively evaluate some speaker-normalization effects of the VTLs is proposed. We employed accumulated means of formant trajectories to estimate the VTLs of speakers ranging from children to adults. For the formant estimation, the inverse-filter control (IFC) system was used. In the system, the decision of analysis order, which means number of formants to be estimated, is automated. Moreover, to evaluate the speaker-normalization effect of VTLs, we proposed the data reduction method, which can reasonably find dense areas of ellipses from distributions in the formant space. Using these ellipse areas, we evaluated the three normalization effects of VTLs: normalization by the mean of all VTLs as the standard, by speaker-categorical means of VTLs, and by individual VTLs. The area reduced from the standard area of the original data by 39.5% and 46.6% in the case of the categorical means and individual VTLs, respectively. As a result, our proposed method was used to provide a “normalized vowel map (NVM)” that visualizes universal vowel-distributions as a core image of linguistic information. Finally, we compared the estimated VTLs with those by another method based on magnetic resonance imaging (MRI) data, using the proposed methods.