Subband based voice conversion

Subband based voice conversion
复制标题

基于子带的语音转换

DOI:
10.21437/icslp.2002-137
复制
发表时间:
2002
影响因子:
5.4
通讯作者:
L. Arslan
L. Arslan
中科院分区:
农林科学1区
文献类型:
--
作者:
O. Türk;L. Arslan

文献摘要

被引文献

相似文献

提出了一种新的语音转换方法,在较高采样率下提高了语音转换输出的质量。对基于分段码本的说话人变换算法(STASC)进行了改进,使其能够处理不同子带上的源和目标语音频谱。新的方法确保了在16 KHz以上的采样率下有更好的转换。采用离散小波变换(DWT)进行子带分解,以更高的分辨率更好地估计语音频谱。由于以较低的采样率降低了计算复杂度,因此实现了更快的语音转换。并利用该算法实现了一个语音转换系统(VCS)。主观听力测试和对电影配音和循环的应用都证明了该方法的有效性。在ABX听力测试中,与基于全频带的输出相比,听者更喜欢基于子带的输出92.1%。
A new voice conversion method that improves the quality of the voice conversion output at higher sampling rates is proposed. Speaker Transformation Algorithm Using Segmental Codebooks (STASC) is modified to process source and target speech spectra in different subbands. The new method ensures better conversion at sampling rates above 16KHz. Discrete Wavelet Transform (DWT) is employed for subband decomposition to estimate the speech spectrum better with higher resolution. Faster voice conversion is achieved since the computational complexity decreases at a lower sampling rate. A Voice Conversion System (VCS) is implemented using the proposed algorithm with necessary tools. The performance of the proposed method is demonstrated by both subjective listening tests and applications to film dubbing and looping. In ABX listening tests, the listeners preferred the subband based output by 92.1% as compared to the full-band based output.