SABR: sparse, anchor-based representation of the speech signal

SABR: sparse, anchor-based representation of the speech signal
复制标题

SABR:语音信号的稀疏、基于锚的表示

DOI:
--
复制
发表时间:
2015
期刊:
Interspeech
影响因子:
--
通讯作者:
R. Gutierrez
R. Gutierrez
中科院分区:
--
文献类型:
--
作者:
C. Liberatore;S. Aryal;Zelun Wang;Seth Polsley;R. Gutierrez

文献摘要

被引文献

相似文献

我们提出了SABR(稀疏,基于锚点的表示),分析技术将语音信号分解成说话人相关和说话人无关的组件。给定特定说话者的话语集合,SABR使用每个音素的质心作为声学“锚”,然后应用Lasso正则化来将每个语音帧表示为锚的稀疏非负组合。我们说明了一个独立于说话人的音素识别任务和语音转换任务的方法的性能。使用线性分类器,SABR权重实现显着更高的音素识别率比梅尔频率倒谱系数。SABR权重也可以直接用于执行口音转换,而无需训练说话者到说话者回归模型。
We present SABR (Sparse, Anchor-Based Representation), an analysis technique to decompose the speech signal into speaker-dependent and speaker-independent components. Given a collection of utterances for a particular speaker, SABR uses the centroid for each phoneme as an acoustic “anchor,” then applies Lasso regularization to represent each speech frame as a sparse non-negative combination of the anchors. We illustrate the performance of the method on a speaker-independent phoneme recognition task and a voice conversion task. Using a linear classifier, SABR weights achieve significantly higher phoneme recognition rates than Mel frequency Cepstral coefficients. SABR weights can also be used directly to perform accent conversion without the need to train a speakerto-speaker regression model.