Speaker Adaptation on Articulation and Acoustics for Articulation-to-Speech Synthesis.

Speaker Adaptation on Articulation and Acoustics for Articulation-to-Speech Synthesis.
复制标题

DOI:
10.3390/s22166056
复制
发表时间:
2022-08-13
期刊:
Sensors (Basel, Switzerland)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

无声语音接口(SSI)将非音频生物信号(如发音运动)转换为语音。这项技术有可能恢复失去声音但仍能发音的人的言语能力(例如,喉切除者)。语音合成(ATS)是一种语音合成算法设计,具有易于实现和低延迟的优点,因此越来越受到人们的青睐。目前的ATS研究集中在说话者相关(SD)模型,以避免发音模式和声学特征在说话者之间的大的变化。然而,这些设计受到来自单个扬声器的小数据大小的限制。包括多个扬声器的数据的扬声器自适应设计有可能解决单个扬声器的数据大小有限的问题,但是,很少有以前的研究已经调查了它们在ATS中的性能。在本文中,我们使用公开可用的电磁发音(EMA)数据集研究了输入发音和输出声学信号(无论是否直接包含来自测试扬声器的数据)上的扬声器自适应。我们使用Procrustes匹配和语音转换的清晰度和语音适应,分别。ATS模型的性能通过梅尔倒谱失真(MCDs)客观地测量。合成语音样本已生成,并在补充材料中提供。实验结果表明,Procrustes匹配和语音转换对非特定人自动测试系统的性能有明显的改善。在训练过程中直接包括目标说话人数据,说话人自适应ATS实现了与说话人相关ATS相当的性能。据我们所知,这是第一个研究表明,扬声器自适应ATS可以实现非统计上不同的性能扬声器依赖ATS。
Silent speech interfaces (SSIs) convert non-audio bio-signals, such as articulatory movement, to speech. This technology has the potential to recover the speech ability of individuals who have lost their voice but can still articulate (e.g., laryngectomees). Articulation-to-speech (ATS) synthesis is an algorithm design of SSI that has the advantages of easy-implementation and low-latency, and therefore is becoming more popular. Current ATS studies focus on speaker-dependent (SD) models to avoid large variations of articulatory patterns and acoustic features across speakers. However, these designs are limited by the small data size from individual speakers. Speaker adaptation designs that include multiple speakers’ data have the potential to address the issue of limited data size from single speakers; however, few prior studies have investigated their performance in ATS. In this paper, we investigated speaker adaptation on both the input articulation and the output acoustic signals (with or without direct inclusion of data from test speakers) using the publicly available electromagnetic articulatory (EMA) dataset. We used Procrustes matching and voice conversion for articulation and voice adaptation, respectively. The performance of the ATS models was measured objectively by the mel-cepstral distortions (MCDs). The synthetic speech samples were generated and are provided in the supplementary material. The results demonstrated the improvement brought by both Procrustes matching and voice conversion on speaker-independent ATS. With the direct inclusion of target speaker data in the training process, the speaker-adaptive ATS achieved a comparable performance to speaker-dependent ATS. To our knowledge, this is the first study that has demonstrated that speaker-adaptive ATS can achieve a non-statistically different performance to speaker-dependent ATS.
DOI: 10.1007/bf02291478
发表时间: 1975-01-01
期刊: PSYCHOMETRIKA
影响因子: 3
作者:
GOWER, JC
通讯作者: GOWER, JC
DOI: 10.2147/mder.s133225
发表时间: 2017
期刊: Medical devices (Auckland, N.Z.)
影响因子: --
作者:
Kaye R;Tang CG;Sinclair CF
通讯作者: Sinclair CF
DOI: 10.1109/taslp.2017.2757263
发表时间: 2017-12-01
影响因子: 5.4
作者:
Gonzalez, Jose A.;Cheah, Lam A.;Holdsworth, Ed
通讯作者: Holdsworth, Ed
DOI: 10.1590/s1807-59322005000200010
发表时间: 2005-04-01
期刊: Clinics
影响因子: 2.7
作者:
Braz, Daniella Scalet Amorin;Ribas, Marta Maria;Barros, Ana Paula Brandão
通讯作者: Barros, Ana Paula Brandão
DOI: 10.1016/j.specom.2009.11.004
发表时间: 2010-04-01
影响因子: 3.2
作者:
Hueber, Thomas;Benaroya, Elie-Laurent;Stone, Maureen
通讯作者: Stone, Maureen