Augmented speech communication using multi-modal signals with real-time, low-latency voice conversion
Augmented speech communication using multi-modal signals with real-time, low-latency voice conversion
批准号:
22KJ1519
负责人:
HUANG WENCHIN
金额:
$1.41万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for JSPS Fellows
财政年份:
2023
资助国家:
日本
项目状态:
已结题
起止时间:
2023-03-08 至 2024-03-31
中文摘要
本研究的目的是应用语音转换(VC)技术,借助多模式信号和实时处理技术,实现一种面向现实应用的交互式语音生成范例。在第二年,申请人专注于三个方面:(1)继续改进基本的VC技术,特别是基于自监督语音表示(S3R)的VC,这是一个减少训练数据需求的新兴趋势。申请者不断更新S3PRL-VC,这是一个为研究人员评估VC的S3R模型的开源工具包,并在IEEE信号处理精选主题杂志上发表了最新的实验结果。(2)外国口音转换,这是一项有助于减少外国口音以实现有效沟通的任务。一篇对当前方法提供统一评估并确定未解决问题的论文被提交给一个国际会议,目前正在进行审查。(3)歌唱声音转换,这是一种有可能增强人类交流能力的基本技术。申请者正在开展一项名为2023歌声转换挑战的科学活动,旨在提供一个包括任务和数据集的统一实验环境,以吸引世界各地的研究人员研究这一问题,并探索最先进技术的局限性。
英文摘要
The purpose of this research is to apply voice conversion (VC) to realize an interactive speech production paradigm for real-world applications, with the help of multimodal signals and real-time processing techniques. In the second year, the applicant focused on three aspects.(1) Continued improvement on fundamental VC techniques, specifically self-supervised speech representation (S3R)-based VC, an emerging trend which reduces training data requirements. The applicant kept on updating S3PRL-VC, an open-source toolkit for researchers to evaluate S3R models for VC, and published the latest experimental results in the IEEE Journal of Selected Topics in Signal Processing.(2) Foreign accent conversion, a task that helps reduce foreign accents for efficient communication. A paper that provides an unified evaluation of current approaches and identifies unsolved problems is submitted to an international conference and currently under review.(3) Singing voice conversion, a fundamental technique that has the potential to augment the communication ability of human. The applicant is running a scientific event named the Singing Voice Conversion Challenge 2023, which aims to provide an unified experimental setting including task and dataset, in order to attract researchers world-wide to look into this problem and explore the limitation of the state-of-the-art techniques.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.21437/interspeech.2021-208
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
作者:
[Wen-Chin Huang;Kazuhiro Kobayashi;Yu-Huai Peng;Ching-Feng Liu;Yu Tsao;Hsin-Min Wang;T. Toda]
通讯作者:
Wen-Chin Huang;Kazuhiro Kobayashi;Yu-Huai Peng;Ching-Feng Liu;Yu Tsao;Hsin-Min Wang;T. Toda
CRANK: an Open-Source Software for Nonparallel Voice Conversion based on Vetor-Quantized Variational Autoencoder
CRANK:基于矢量量化变分自动编码器的非并行语音转换开源软件
DOI:
--
发表时间:
2021
期刊:
影响因子:
--
作者:
[Kazuhiro Kobayashi, Wen-Chin Huang, Yi-Chiao Wu, Patrick Tobing, Tomoki Hayashi, and Tomoki Toda]
通讯作者:
and Tomoki Toda
DOI:
10.1109/asru51503.2021.9688010
发表时间:
2021-07
期刊:
2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
影响因子:
--
作者:
[Wen-Chin Huang;Tomoki Hayashi;Xinjian Li;Shinji Watanabe;T. Toda]
通讯作者:
Wen-Chin Huang;Tomoki Hayashi;Xinjian Li;Shinji Watanabe;T. Toda
DOI:
10.1109/icassp43922.2022.9746430
发表时间:
2021-10
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[Wen-Chin Huang;Shu-Wen Yang;Tomoki Hayashi;Hung-yi Lee;Shinji Watanabe;T. Toda]
通讯作者:
Wen-Chin Huang;Shu-Wen Yang;Tomoki Hayashi;Hung-yi Lee;Shinji Watanabe;T. Toda
海外基金