SABR: sparse, anchor-based representation of the speech signal
SABR: sparse, anchor-based representation of the speech signal
复制标题
SABR:语音信号的稀疏、基于锚的表示
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
R. Gutierrez
中科院分区:
文献类型:
--
作者:
C. Liberatore;S. Aryal;Zelun Wang;Seth Polsley;R. Gutierrez
We present SABR (Sparse, Anchor-Based Representation), an analysis technique to decompose the speech signal into speaker-dependent and speaker-independent components. Given a collection of utterances for a particular speaker, SABR uses the centroid for each phoneme as an acoustic “anchor,” then applies Lasso regularization to represent each speech frame as a sparse non-negative combination of the anchors. We illustrate the performance of the method on a speaker-independent phoneme recognition task and a voice conversion task. Using a linear classifier, SABR weights achieve significantly higher phoneme recognition rates than Mel frequency Cepstral coefficients. SABR weights can also be used directly to perform accent conversion without the need to train a speakerto-speaker regression model.