Automatic discovery of a phonetic inventory for unwritten languages for statistical speech synthesis

Automatic discovery of a phonetic inventory for unwritten languages for statistical speech synthesis
复制标题

自动发现非书面语言的语音库存以进行统计语音合成

DOI:
10.1109/icassp.2014.6854069
复制
发表时间:
2014
期刊:
2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
A. Black
A. Black
中科院分区:
--
文献类型:
--
作者:
P. Muthukumar;A. Black

文献摘要

被引文献

相似文献

语音合成系统通常是用语音数据和转录构建的。在本文中,我们试图在没有转录或语言知识可用的情况下构建合成系统。通常至少掌握一门语言的语音知识是必要的。本文提出了一种利用发音特征(articatory Features, AFs)自动获取语料库语音知识的方法。使用三隐层神经网络在任意其他语言的自举语料库上训练发音特征预测器。该神经网络在语音语料库上运行,以提取语音特征。分层聚类是用来将AFs聚类为类别,即电话。通过计算每个集群中af的平均值来获得每个推断电话的语音信息。报告了使用该框架以多种语言构建的系统的结果。
Speech synthesis systems are typically built with speech data and transcriptions. In this paper, we try to build synthesis systems when no transcriptions or knowledge about the language are available. It is usually necessary to at least possess phonetic knowledge about the language. In this paper, we propose an automated way of obtaining phones and phonetic knowledge about the corpus at hand by making use of Articulatory Features (AFs). An Articulatory Feature predictor is trained on a bootstrap corpus in an arbitrary other language using a three-hidden layer neural network. This neural network is run on the speech corpus to extract AFs. Hierarchical clustering is used to cluster the AFs into categories i.e. phones. Phonetic information about each of these inferred phones is obtained by computing the mean of the AFs in each cluster. Results of systems built with this framework in multiple languages are reported.