Using articulatory features and inferred phonological segments in zero resource speech processing

Using articulatory features and inferred phonological segments in zero resource speech processing
复制标题

在零资源语音处理中使用发音特征和推断的语音片段

DOI:
10.21437/interspeech.2015-643
复制
发表时间:
2015
期刊:
2008 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
A. Black
A. Black
中科院分区:
--
文献类型:
--
作者:
P. Baljekar;Sunayana Sitaram;P. Muthukumar;A. Black

文献摘要

被引文献

相似文献

无监督的子词单元发现是零资源语言识别和合成中的一个重要问题,在零资源语言中,音素集可能是未知的,唯一可用的资源是语音。我们使用的技术,我们最近开发的非常低的资源语言,没有书面形式来发现这样的单位建立合成的声音。我们使用在更高资源语言中的标记语音上训练的发音特征来推断不同粒度的语音段。我们使用原始的发音特征和推断单元的发音特征作为语音的基于帧的表示。我们评估我们的技术最小对ABX歧视内和跨扬声器。此外,利用持续时间的信息,我们从推断的语音单位,我们还提出了梅尔倒谱系数失真,语音合成质量的客观度量的评价结果。我们评估我们的技术在多个数据库的英语,也对特松加语和印度语,我们在其中应用上述方法跨语言。
Unsupervised discovery of subword units is an important problem in recognition and synthesis of zero-resource languages, in which phonesets may not be known and the only resource that may be available is speech. We use techniques that we have recently developed for building synthetic voices for very low resource languages without a written form to discover such units. We use Articulatory Features trained on labeled speech in a higher resource language to infer phonological segments of varying granularity. We use both the raw Articulatory Features and the Articulatory Features of the inferred units as frame-based representations of speech. We evaluate our techniques on minimal pair ABX discrimination within and across speakers. In addition, to exploit the duration information we get from the inferred phonological units, we also present evaluation results on Mel Cepstral Distortion, an objective metric of speech synthesis quality. We evaluate our techniques on multiple databases of English, and also on Tsonga and Indic languages, in which we apply the above methods cross-lingually.