An Efficient Image to Sound Mapping Method Using Speech Spectral Phase and Multi-Column Image

An Efficient Image to Sound Mapping Method Using Speech Spectral Phase and Multi-Column Image
复制标题

DOI:
10.1587/transfun.e100.a.893
复制
发表时间:
2017-03-01
影响因子:
0.5
通讯作者:
Iiguni, Youji
Iiguni, Youji
中科院分区:
计算机科学4区
文献类型:
--
作者:
Kawamura, Arata;Igarashi, Hiro;Iiguni, Youji

文献摘要

被引文献

相似文献

图像到声音映射是一种将图像转换为声音信号的技术,该信号随后被视为声音频谱图。一般来说,变换后的声音不同于人类语音信号。本文提出了一种有效的图像到声音映射方法,该方法无需任何训练即可提供可理解的语音信号。为了合成这样的语音信号,所提出的方法利用多列图像和从对语音的长时间观察中获得的语​​音频谱相位。可以从合成语音信号的声谱图中检索原始图像。使用客观测试来评估合成的语音和重建的图像质量。
Image-to-sound mapping is a technique that transforms an image to a sound signal, which is subsequently treated as a sound spectrogram. In general, the transformed sound differs from a human speech signal. Herein an efficient image-to-sound mapping method, which provides an understandable speech signal without any training, is proposed. To synthesize such a speech signal, the proposed method utilizes a multi-column image and a speech spectral phase that is obtained from a long-time observation of the speech. The original image can be retrieved from the sound spectrogram of the synthesized speech signal. The synthesized speech and the reconstructed image qualities are evaluated using objective tests.