Multimodal modeling and validation of simplified vocal tract acoustics for sibilant /s/

Multimodal modeling and validation of simplified vocal tract acoustics for sibilant /s/
复制标题

咝咝音 /s/ 简化声道声学的多模态建模和验证

DOI:
10.1016/j.jsv.2017.09.004
复制
发表时间:
2017
影响因子:
4.7
通讯作者:
Wada S.
Wada S.
中科院分区:
工程技术2区
文献类型:
--
作者:
Yoshinaga T.;Van Hirtum A.;Wada S.

文献摘要

相似文献

为了研究咝音 /s/ 的声学特性,多模态理论被应用于简化的声道几何形状,该几何形状是通过收集声谱的单个说话人的 CT 扫描得出的。声道由具有矩形横截面和恒定宽度的波导串联表示,声源放置在声道的入口处或代表咝擦音槽的收缩处的下游。使用声学驱动器或声道入口处的气流供应对建模的压力幅度进行了实验验证。结果表明,利用入口处的声源预测的频谱(包括高阶模式)与利用入口处的声驱动器测量的频谱相匹配。使用颈缩下游的源建模的频谱捕获了在 4 kHz 处观察到的扬声器的第一个特征峰值。通过将源放置在上牙壁附近,可以预测扬声器在 8 kHz 处观察到的较高频率峰值,其中包含高阶模式。当声源放置在缩窄处下游时,在简化声道中观察到压力振幅的特征峰频率、节点和波腹。这些结果表明,多模态方法能够捕获频谱中峰值的幅度和频率以及由于声道内部 /s/ 造成的压力分布的节点和波腹。
To investigate the acoustic characteristics of sibilant /s/, multimodal theory is applied to a simplified vocal tract geometry derived from a CT scan of a single speaker for whom the sound spectrum was gathered. The vocal tract was represented by a concatenation of waveguides with rectangular cross-sections and constant width, and a sound source was placed either at the inlet of the vocal tract or downstream from the constriction representing the sibilant groove. The modeled pressure amplitude was validated experimentally using an acoustic driver or airflow supply at the vocal tract inlet. Results showed that the spectrum predicted with the source at the inlet and including higher-order modes matched the spectrum measured with the acoustic driver at the inlet. Spectra modeled with the source downstream from the constriction captured the first characteristic peak observed for the speaker at 4 kHz. By positioning the source near the upper teeth wall, the higher frequency peak observed for the speaker at 8 kHz was predicted with the inclusion of higher-order modes. At the frequencies of the characteristic peaks, nodes and antinodes of the pressure amplitude were observed in the simplified vocal tract when the source was placed downstream from the constriction. These results indicate that the multimodal approach enables to capture the amplitude and frequency of the peaks in the spectrum as well as the nodes and antinodes of the pressure distribution due to /s/ inside the vocal tract.