Speech transformations based on a sinusoidal representation
Speech transformations based on a sinusoidal representation
复制标题
DOI:
10.1109/icassp.1985.1168380
复制
发表时间:
1985-04
期刊:
影响因子:
--
通讯作者:
T. Quatieri;R. McAulay
中科院分区:
文献类型:
--
作者:
T. Quatieri;R. McAulay
This paper presents a new speech analysis/synthesis technique based on a sinusoidal representation of the speech production mechanism but which is independent of pitch and the voiced/unvoiced speech state. The resulting synthetic speech preserves the waveform shape and is essentially perceptually indistinguishable from the original. The method provides the basis for a general class of speech transformations and is successfully applied to time-scale modification, frequency scaling, and scaling of pitch. Furthermore, these modifications can be performed with a time-varying rate of change, allowing, for example, continuous adjustment of a speaker's fundamental frequency and rate of articulation. Although the analysis/synthesis system was originally designed for single-speaker signals, it is equally capable of recovering and modifying nonspeech signals such as music, multi-speakers, marine biologic sounds, and speech in the presence of interferences such as noise and musical backgrounds.