VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop

VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop
复制标题

DOI:
--
复制
发表时间:
2017-07
期刊:
arXiv: Learning
影响因子:
--
通讯作者:
Yaniv Taigman;Lior Wolf;Adam Polyak;Eliya Nachmani
Yaniv Taigman;Lior Wolf;Adam Polyak;Eliya Nachmani
中科院分区:
其他
文献类型:
--
作者:
Yaniv Taigman;Lior Wolf;Adam Polyak;Eliya Nachmani

文献摘要

被引文献

相似文献

我们提出了一种新的神经文本到语音(TTS)的方法,能够将文本转换为语音在野外采样的声音。与其他系统不同,我们的解决方案能够处理不受约束的语音样本,而不需要对齐的音素或语言特征。该网络结构比现有文献中的结构简单,基于一种新型的移位缓冲工作记忆。相同的缓冲器用于估计注意力、计算输出音频以及用于更新缓冲器本身。输入的句子使用上下文无关的查找表进行编码,该查找表包含每个字符或音素的一个条目。说话者同样由一个短向量表示,即使只有几个样本,也可以拟合新的身份。通过在生成音频之前启动缓冲器来实现所生成的语音的可变性。在多个数据集上的实验结果证明了令人信服的能力,使得TTS可用于更广泛的应用。为了提高可重复性,我们发布了我们的源代码和模型。
We present a new neural text to speech (TTS) method that is able to transform text to speech in voices that are sampled in the wild. Unlike other systems, our solution is able to deal with unconstrained voice samples and without requiring aligned phonemes or linguistic features. The network architecture is simpler than those in the existing literature and is based on a novel shifting buffer working memory. The same buffer is used for estimating the attention, computing the output audio, and for updating the buffer itself. The input sentence is encoded using a context-free lookup table that contains one entry per character or phoneme. The speakers are similarly represented by a short vector that can also be fitted to new identities, even with only a few samples. Variability in the generated speech is achieved by priming the buffer prior to generating the audio. Experimental results on several datasets demonstrate convincing capabilities, making TTS accessible to a wider range of applications. In order to promote reproducibility, we release our source code and models.