Speech transformations based on a sinusoidal representation

Speech transformations based on a sinusoidal representation
复制标题

DOI:
10.1109/icassp.1985.1168380
复制
发表时间:
1985-04
期刊:
ICASSP '85. IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
T. Quatieri;R. McAulay
T. Quatieri;R. McAulay
中科院分区:
其他
文献类型:
--
作者:
T. Quatieri;R. McAulay

文献摘要

被引文献

相似文献

本文提出了一种新的语音分析/合成技术的基础上的正弦表示的语音产生机制,但它是独立的音高和有声/清音语音状态。所得到的合成语音保留了波形形状,并且基本上在感知上与原始语音无法区分。该方法提供了一个通用类的语音变换的基础上,并成功地应用于时间尺度的修改,频率缩放,和缩放的音调。此外,这些修改可以以随时间变化的变化率来执行,从而允许例如连续调整说话者的基频和发音率。虽然分析/合成系统最初是为单扬声器信号设计的,但它同样能够在存在噪声和音乐背景等干扰的情况下恢复和修改非语音信号,如音乐、多扬声器、海洋生物声音和语音。
This paper presents a new speech analysis/synthesis technique based on a sinusoidal representation of the speech production mechanism but which is independent of pitch and the voiced/unvoiced speech state. The resulting synthetic speech preserves the waveform shape and is essentially perceptually indistinguishable from the original. The method provides the basis for a general class of speech transformations and is successfully applied to time-scale modification, frequency scaling, and scaling of pitch. Furthermore, these modifications can be performed with a time-varying rate of change, allowing, for example, continuous adjustment of a speaker's fundamental frequency and rate of articulation. Although the analysis/synthesis system was originally designed for single-speaker signals, it is equally capable of recovering and modifying nonspeech signals such as music, multi-speakers, marine biologic sounds, and speech in the presence of interferences such as noise and musical backgrounds.