RI: Small: Creating Text-to-Speech Synthesis for Low Resource Languages
RI: Small: Creating Text-to-Speech Synthesis for Low Resource Languages
批准号:
1717680
负责人:
Julia Hirschberg
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2021-08-31
中文摘要
最近语音技术的进步导致了语音对话系统(SDS)的广泛使用,例如Siri (iPhone)和b谷歌Assistant (Android)。这些系统支持通过语音对英语、法语、普通话、日语和西班牙语等高级资源语言的信息访问进行重大改进。对于这些语言,研究人员已经建立了字典、解析器、词性标注器、语言模型、搜索引擎和机器翻译引擎来支持语音技术。然而,世界上大约有6500种语言,包括他加禄语、泰米尔语、斯瓦希里语、越南语和普什图语,其中许多语言被数百万人使用,但它们没有获得构建SDS所需的计算资源。这些被称为低资源语言(LRLs)。lrl的讲话者没有从hrl的讲话者所拥有的相同的通信和搜索功能中受益。特别是,很少有研究和资源支持开发文本到语音合成(TTS)系统,以在这些语言中为SDS生成类似siri的语音。此外,商业和研究TTS系统还需要大量仔细记录的单扬声器语音数据,这为LRLs开发TTS创造了另一个主要(且昂贵)的障碍。这项工作将在LRLs中创建TTS系统,并在此过程中为其他人创建和提供工具,以使用“发现”数据(为其他目的记录的数据或在网络上可用的数据)创建他们自己的系统。TTS合成(参数合成和深度神经网络的使用)的新范例正在开发中,这使得理论上可以快速廉价地构建系统,而不需要记录大型专用语音语料库,而是使用为其他目的(如训练语音识别器)记录的数据。这项工作将研究使用这些技术来生产LRL的TTS系统。将探讨两个主要问题:1)过滤发现的数据(例如,去除太大声、太嘈杂或不流畅的数据)以获得可理解且听起来自然的结果的最佳技术是什么?2)使用众包和为语音识别工具,是否可以识别语音识别语言的基本韵律特征,如短语和重音?对英语的初步研究表明,通过使用根据音高变化和发音水平等特征选择的数据子集,可以创建更自然、更容易理解的声音。这些方法将在土耳其语、阿姆哈拉语和泰卢固语等语言中进行测试。评估将根据可理解性和自然度自动进行,并使用众包技术与每种语言的母语人士进行评估。这项探索性工作的最终目标将是在各种各样的LRLs上测试这些技术,这些LRLs已被收集用于开发语音识别器。
英文摘要
Recent advances in speech technology have resulted in wide use of Spoken Dialogue Systems (SDS) such as Siri (iPhone) and Google Assistant (Android). These systems support major improvements in information access by voice for High Resource Languages (HRLs) such as English, French, Mandarin, Japanese, and Spanish. For these languages, researchers have built dictionaries, parsers, part-of-speech taggers, language models, search engines, and machine translation engines to support speech technologies. However, there are ~6500 world languages, including Tagalog, Tamil, Swahili, Vietnamese and Pashto, many of which are spoken by millions of people, but which do not enjoy the computational resources necessary to build SDS. These are termed Low Resource Languages (LRLs). Speakers of LRLs do not benefit from the same communication and search capabilities speakers of HRLs do. In particular, there is little research and few resources supporting the development of Text-to-Speech Synthesis (TTS) systems to produce Siri-like speech for SDS in these languages. Furthermore, both commercial and research TTS systems also require large amounts of carefully recorded, single-speaker speech data, creating another major (and expensive) barrier to TTS development for LRLs. This work will create TTS systems in LRLs and, in the process, create and make available tools for others to create their own systems using "found" data - data recorded for other purposes or available on the web.New paradigms for TTS synthesis (parametric synthesis and the use of Deep Neural Nets) are now being developed which make it theoretically possible to build systems quickly and cheaply without recording large, special-purpose speech corpora, instead using data recorded for other purposes such as training speech recognizers. This work will investigate the use these techniques to produce TTS systems for LRL. Two major problems will be explored: 1) What are the best techniques to filter found data (removing data that is too loud, too noisy or disfluent, for example) to obtain intelligible and natural-sounding results? 2) Can basic prosodic features of LRLs such as phrasing and emphasis be identified, using crowdsourcing and tools developed for HRLs? Pilot studies on English have revealed that more natural and intelligible voices can be created by using subsets of the data selected on features such as pitch variation and level of articulation. These methods will be tested on LRLs such as Turkish, Amharic, and Telugu. Evaluations will be made in terms of intelligibility and naturalness both automatically and using crowdsourcing techniques with native speakers of each language. The ultimate goal of this exploratory work will be to test these techniques on a broad variety of LRLs which have been collected for purposes of developing speech recognizers.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.21437/speechprosody.2018-140
发表时间:
2018-06
期刊:
Speech Prosody 2018
影响因子:
--
作者:
[Erica Cooper;E. Li;Julia Hirschberg]
通讯作者:
Erica Cooper;E. Li;Julia Hirschberg
Adaptation and Frontend Features to Improve Naturalness in Found-Data Synthesis
适应和前端功能可提高发现数据合成的自然度
DOI:
10.21437/speechprosody.2018-160
发表时间:
2018
期刊:
Speech Prosody 2018
影响因子:
--
作者:
[Cooper, Erica, Hirschberg, Julia]
通讯作者:
Hirschberg, Julia
DOI:
10.21437/ssw.2019-48
发表时间:
2019-09
期刊:
10th ISCA Workshop on Speech Synthesis (SSW 10)
影响因子:
--
作者:
[Rose Sloan;S. S. Akhtar-S.;Bryan Li;Ritvik Shrivastava;Agustin Gravano;Julia Hirschberg]
通讯作者:
Rose Sloan;S. S. Akhtar-S.;Bryan Li;Ritvik Shrivastava;Agustin Gravano;Julia Hirschberg
A Comparison of Speaker-based and Utterance-based Data Selection for Text-to-Speech Synthesis
文本转语音合成中基于说话者和基于话语的数据选择的比较
DOI:
10.21437/interspeech.2018-1313
发表时间:
2018
期刊:
Interspeech 2018
影响因子:
--
作者:
[Kai-Zhan Lee, Erica Cooper]
通讯作者:
Kai-Zhan Lee, Erica Cooper
Subset Selection, Adaptation, Gemination and Prosody Prediction for Amharic Text-to-Speech Synthesis
阿姆哈拉语文本转语音合成的子集选择、适应、双生和韵律预测
DOI:
10.21437/ssw.2019-37
发表时间:
2019
期刊:
10th ISCA Speech Synthesis Workshop
影响因子:
--
作者:
[Tesfaye Biru, Elshadai, Tofik Mohammed, Yishak, Tofu, David, Cooper, Erica, Hirschberg, Julia]
通讯作者:
Hirschberg, Julia
共 6 条
EAGER: Identifying and Producing Code-Switching in Languages from Spoken, Lexical and Socio-linguistic Features
-
批准号:2327564
-
项目类别:Standard Grant
-
资助金额:$10.89万
-
财政年份:2023
-
负责人:Julia Hirschberg
-
依托单位:
EAGER: Creating Speech Synthesizers for Low Resource Languages
-
批准号:1548092
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Julia Hirschberg
-
依托单位:
Using Computational Tools to Facilitate Corpus Collection and Language Use in Arrernte (aer)
-
批准号:1160700
-
项目类别:Standard Grant
-
资助金额:$9.82万
-
财政年份:2012
-
负责人:Julia Hirschberg
-
依托单位:
Collaborative Research: CI-P: Reciprosody - A Repository for Prosodically Annotated Material
-
批准号:1205450
-
项目类别:Standard Grant
-
资助金额:$2.5万
-
财政年份:2012
-
负责人:Julia Hirschberg
-
依托单位:
IGERT: From Data to Solutions: A New PhD Program in Transformational Data & Information Sciences Research and Innovation
-
批准号:1144854
-
项目类别:Continuing Grant
-
资助金额:$300.0万
-
财政年份:2012
-
负责人:Julia Hirschberg
-
依托单位:
CI-P: Collaborative Research: Summarizing Opinion and Speaker Attitude in Speech
-
批准号:1059260
-
项目类别:Standard Grant
-
资助金额:$3.79万
-
财政年份:2011
-
负责人:Julia Hirschberg
-
依托单位:
EAGER: Using Social Media and Crowdsourcing to Create a New Affect Dictionary
-
批准号:1145505
-
项目类别:Standard Grant
-
资助金额:$0.62万
-
财政年份:2011
-
负责人:Julia Hirschberg
-
依托单位:
RI: Medium: Collaborative Research: From Text to Pictures
-
批准号:0904361
-
项目类别:Standard Grant
-
资助金额:$83.52万
-
财政年份:2009
-
负责人:Julia Hirschberg
-
依托单位:
RI-Medium: Collaborative: Corpus-Based Studies of Lexical, Acoustic-Prosodic, and Discourse Entrainment in Spoken Dialogue
-
批准号:0803148
-
项目类别:Standard Grant
-
资助金额:$44.19万
-
财政年份:2008
-
负责人:Julia Hirschberg
-
依托单位:
Doctoral Consortium at The Human Language Technology Conference - North American chapter of the Association for Computational Linguistics annual meeting (NAACL HLT) 2007.
-
批准号:0707305
-
项目类别:Standard Grant
-
资助金额:$1.92万
-
财政年份:2007
-
负责人:Julia Hirschberg
-
依托单位:
Collaborative Research: Translating Prosody in an English/Chinese Language Tutoring System
-
批准号:0534568
-
项目类别:Continuing Grant
-
资助金额:$31.07万
-
财政年份:2006
-
负责人:Julia Hirschberg
-
依托单位:
ITR: Recognizing and Understanding Emotion in Speech
-
批准号:0325399
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Julia Hirschberg
-
依托单位:
Collaborative Research: Monitoring Student State in Tutorial Spoken Dialogue
-
批准号:0328295
-
项目类别:Continuing Grant
-
资助金额:$27.0万
-
财政年份:2003
-
负责人:Julia Hirschberg
-
依托单位:
Dialogue Prosody in Interactive Voice Response Systems
-
批准号:0307905
-
项目类别:Continuing Grant
-
资助金额:$53.06万
-
财政年份:2003
-
负责人:Julia Hirschberg
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: