Language-independent, multi-modal, and data-efficient approaches for speech synthesis and translation
Language-independent, multi-modal, and data-efficient approaches for speech synthesis and translation
批准号:
21K11951
负责人:
Cooper Erica
金额:
$2.66万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2021
资助国家:
日本
项目状态:
已结题
起止时间:
2021-04-01 至 2024-03-31
中文摘要
在该项目的第二年,我们考察了两个主要主题:使用自我监督语音表征的低资源语言的独立于语言的、数据高效的文本到语音合成,以及自动平均评分预测。自我监督语音表征在许多下游语音相关任务中显示出显著的实用性,并且已被证明包含语音信息。因此,我们选择这些作为文本到语音合成的中间表示,基于来自许多语言的数据进行训练,然后只使用少量数据就可以微调到一种新的语言。这项工作正在进行中,我们正在与加拿大国家研究委员会和爱丁堡大学的研究人员合作。我们还将合成语音的自动评估确定为低资源语言的一个重要主题,因为对于这些语言来说,找到参与听力测试的听者可能特别困难。我们与名古屋大学和中央研究院合作,共同举办了第一届VoiceMOS挑战赛,这是一项针对合成语音的自动平均意见得分(MOS)预测的共同任务。这项挑战吸引了来自学术界和产业界的22支参赛队伍,我们在InterSpeech 2022上举办了一场关于这项挑战的特别会议。这一挑战引起了人们对这个话题的极大兴趣,从而推动了这一领域的发展。
英文摘要
In this second year of the project, we looked at two main topics: language-independent, data-efficient text-to-speech synthesis for low-resource languages using self-supervised speech representations, and automatic mean opinion score prediction.Self-supervised representations for speech have shown remarkable usefulness for many downstream speech-related tasks, and have been shown to contain phonetic information. We therefore chose these as an intermediate representation for text-to-speech synthesis trained on data from many languages, which can then be fine-tuned to a new language using only a small amount of data. This is ongoing work in progress, and we are collaborating with researchers from the National Research Council of Canada and the University of Edinburgh.We have also identified automatic evaluation of synthesized speech as an important topic for low-resource languages, since finding listeners to participate in listening tests can be especially difficult for these languages. In collaboration with Nagoya University and Academia Sinica, we co-organized the first VoiceMOS Challenge, a shared task for automatic mean opinion score (MOS) prediction for synthesized speech. The challenge attracted 22 participating teams from academia and industry, and we ran a special session about the challenge at Interspeech 2022. This challenge has advanced the field by generating a great deal of interest in this topic.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
University of Edinburgh(英国)
爱丁堡大学(英国)
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
DOI:
10.1109/icassp43922.2022.9747728
发表时间:
2021-10
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[Cheng-I Lai;Erica Cooper;Yang Zhang;Shiyu Chang;Kaizhi Qian;Yiyuan Liao;Yung-Sung Chuang;Alexander H. Liu;J. Yamagishi;David Cox;James R. Glass]
通讯作者:
Cheng-I Lai;Erica Cooper;Yang Zhang;Shiyu Chang;Kaizhi Qian;Yiyuan Liao;Yung-Sung Chuang;Alexander H. Liu;J. Yamagishi;David Cox;James R. Glass
The VoiceMOS Challenge 2022
2022 年 VoiceMOS 挑战赛
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Wen-Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, Junichi Yamagishi]
通讯作者:
Junichi Yamagishi
The VoiceMOS Challenge 2022 website
VoiceMOS 挑战 2022 网站
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Generalization Ability of MOS Prediction Networks
MOS预测网络的泛化能力
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Erica Cooper, Wen-Chin Huang, Tomoki Toda, Junichi Yamagishi]
通讯作者:
Junichi Yamagishi
共 13 条
Encoder Factorization for Capturing Dialect and Articulation Level in End-to-End Speech Synthesis
-
批准号:19K24372
-
项目类别:Grant-in-Aid for Research Activity Start-up
-
资助金额:$1.83万
-
财政年份:2019
-
负责人:Cooper Erica
-
依托单位:
海外基金