Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism

Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism
复制标题

DOI:
10.1109/icassp49357.2023.10095919
复制
发表时间:
2022-12
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Yukiya Hono;Kei Hashimoto;Yoshihiko Nankaku;K. Tokuda
Yukiya Hono;Kei Hashimoto;Yoshihiko Nankaku;K. Tokuda
中科院分区:
其他
文献类型:
--
作者:
Yukiya Hono;Kei Hashimoto;Yoshihiko Nankaku;K. Tokuda

文献摘要

相似文献

提出了一种基于音符位置感知注意机制的歌唱语音合成模型(Seq2seq)。一种可以同时执行声学和时间建模的seq2seq建模方法很有吸引力。然而,由于歌唱声音的时间建模的困难,目前许多基于编解码器模型的SVS系统仍然明确地依赖于由附加模块产生的持续时间信息。虽然一些研究使用具有注意机制的seq2seq模型进行同步建模,但它们对时间建模的稳健性不够。所提出的注意机制旨在通过考虑乐谱给出的节奏来估计注意权重。此外,还介绍了几种提高歌唱声音建模性能的技术。实验结果表明,该模型在时序的自然性和稳健性方面都是有效的。
This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the temporal modeling of singing voices, many recent SVS systems with an encoder-decoder-based model still rely on explicitly on duration information generated by additional modules. Although some studies perform simultaneous modeling using seq2seq models with an attention mechanism, they have insufficient robustness against temporal modeling. The proposed attention mechanism is designed to estimate the attention weights by considering the rhythm given by the musical score. Furthermore, several techniques are also introduced to improve the modeling performance of the singing voice. Experimental results indicated that the proposed model is effective in terms of both naturalness and robustness of timing.