SBIR Phase II: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
SBIR Phase II: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
批准号:
1555608
负责人:
Rupal Patel
金额:
$74.78万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-04-01 至 2020-09-30
中文摘要
小型企业创新研究(SBIR)第二阶段项目的更广泛影响/商业潜力是为文本到语音应用提供定制的数字语音。我们每个人都有一个独特的声纹--这是我们自我认同的重要组成部分。尽管文语转换技术的质量有所提高,但语音选择仍然有限。对于依赖设备说话的250万美国人(以及全球数千万)来说,获得定制的数字语音是游戏规则的改变者。这是功能性解决方案和以独特身份被倾听之间的区别。更多的社会联系机会提高了生活质量、独立性和获得教育和职业资源的机会,从而缩小了残疾人和非残疾人之间的差距。这种迫在眉睫的社会需求尚未得到满足,再加上与我们交谈和为我们说话的设备的日益普及,为大规模生产的高质量、个性化的数字声音创造了一个引人注目、及时和重要的商业机会。通过利用该公司众包的人工语音库和专有的语音匹配和混合算法,该技术有可能使每个人都能够通过自己的语音表达自己。这个小企业创新研究第二阶段项目建立在该公司NSF资助的研究和第一阶段成果的基础上,支持定制语音构建技术的可行性和商业化。文语转换市场包括辅助技术、企业和消费者应用程序,目前价值约10亿美元,正在迅速增长和成熟,需要创新。为了创建自定义语音,该公司利用了语音产生的源过滤理论。对于那些不能或不愿意录制几个小时的语音的人,该公司会提取一个简短的声音样本--即使是一个元音也包含足够的“声音DNA”来进行个性化处理。然后,信号源的身份线索与该公司Voicebank中人口统计和声学匹配的捐赠者的过滤器属性相结合。其结果是,声音捕捉到接受者的声音身份,但捐赠者的清晰度。第二阶段的技术目标是满足以下需求:1)客户驱动的语音定制;2)众包录音的质量保证;3)语音老化算法;4)定向捐赠者招募算法。这些进展将有助于确保辅助技术的滩头阵地,并刺激虚拟现实、个人机器人和物联网数字角色等更广泛应用的创新。
英文摘要
The broader impact/commercial potential of this Small Business Innovation Research (SBIR) Phase II project is to offer custom crafted digital voices for text-to-speech applications. Each one of us has a unique voiceprint - an essential part of our self-identity. Though the quality of text-to-speech technology has improved, voice options remain limited. For the 2.5 million Americans (and tens of millions worldwide) living with voicelessness who rely on devices to talk, access to a custom digital voice is a game changer. It's the difference between a functional solution and being heard, uniquely, as oneself. Enhanced opportunities for social connection increase quality of life, independence, and access to educational and vocational resources that can narrow the gap between those with and without disability. This immediate unmet societal need, coupled with the increasing proliferation of devices that speak to us and for us, creates a compelling, timely and significant commercial opportunity for high quality, personalized digital voices that can be produced at scale. By leveraging the company's crowdsourced human voicebank and proprietary voice matching and blending algorithms the technology has the potential to empower everyone to express themselves through their own voice.This Small Business Innovation Research Phase II project builds on the company's NSF-funded research and Phase I results that support feasibility and commercialization of a customized voice building technology. The text-to-speech market, encompassing assistive technologies, enterprise and consumer applications, is currently valued at around $1B and is rapidly growing and ripe for innovation. To create custom voices, the company leverages the source-filter theory of speech production. From those who are unable or unwilling to record several hours of speech the company extracts a brief vocal sample - even a single vowel contains enough 'vocal DNA' to seed the personalization process. Identity cues of the source are then combined with filter properties of a demographically and acoustically matched donor in the company's voicebank. The result is a voice that captures the vocal identity of the recipient but the clarity of the donor. Phase II technical objectives address the need for 1) customer-driven voice customization, 2) quality assurance of crowdsourced recordings, 3) voice aging algorithms, and 4) targeted donor recruitment algorithms. These advances will help secure the assistive technology beachhead and spur innovations for broader applications such as virtual reality, personal robotics, and digital persona for the Internet of Things.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SBIR Phase I: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
-
批准号:1447995
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Rupal Patel
-
依托单位:
EAGER: Collaborative Research: Wireless Sensing of Speech Kinematics and Acoustics for Remediation
-
批准号:1449266
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2014
-
负责人:Rupal Patel
-
依托单位:
HCC: Small: Modeling Acoustic and Articulatory Features for Hybrid Synthesis
-
批准号:1116799
-
项目类别:Standard Grant
-
资助金额:$18.09万
-
财政年份:2011
-
负责人:Rupal Patel
-
依托单位:
HCC-Small: Displaying Prosodic Text for Reading Aloud with Expression
-
批准号:0915527
-
项目类别:Continuing Grant
-
资助金额:$49.84万
-
财政年份:2009
-
负责人:Rupal Patel
-
依托单位:
Adapting a Text-to-Speech Synthesizer to Convey User Identity
-
批准号:0712821
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Rupal Patel
-
依托单位:
SGER: Loudmouth - Toward Intelligible Speech Synthesis in Everyday Noise Situations
-
批准号:0509935
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Rupal Patel
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Baryogenesis, Dark Matter and Nanohertz Gravitational Waves from a Dark
Supercooled Phase Transition
-
批准号:24ZR1429700
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:YUICHIRO NAKAI
-
依托单位:
ATLAS实验探测器Phase 2升级
-
批准号:11961141014
-
项目类别:国际(地区)合作与交流项目
-
资助金额:3350万元
-
批准年份:2019
-
负责人:刘衍文
-
依托单位:
地幔含水相Phase E的温度压力稳定区域与晶体结构研究
-
批准号:41802035
-
项目类别:青年科学基金项目
-
资助金额:12.0万元
-
批准年份:2018
-
负责人:张里
-
依托单位:
基于数字增强干涉的Phase-OTDR高灵敏度定量测量技术研究
-
批准号:61675216
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2016
-
负责人:叶青
-
依托单位:
基于Phase-type分布的多状态系统可靠性模型研究
-
批准号:71501183
-
项目类别:青年科学基金项目
-
资助金额:17.4万元
-
批准年份:2015
-
负责人:陈童
-
依托单位:
纳米(I-Phase+α-Mg)准共晶的临界半固态形成条件及生长机制
-
批准号:51201142
-
项目类别:青年科学基金项目
-
资助金额:25.0万元
-
批准年份:2012
-
负责人:张英波
-
依托单位:
连续Phase-Type分布数据拟合方法及其应用研究
-
批准号:11101428
-
项目类别:青年科学基金项目
-
资助金额:23.0万元
-
批准年份:2011
-
负责人:黄卓
-
依托单位:
D-Phase准晶体的电子行为各向异性的研究
-
批准号:19374069
-
项目类别:面上项目
-
资助金额:6.4万元
-
批准年份:1993
-
负责人:张殿琳
-
依托单位: