Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix Factorization

Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix Factorization
复制标题

DOI:
10.1109/icassp49357.2023.10096401
复制
发表时间:
2022-10
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Jiachen Lian;A. Black;Yijingxiu Lu;L. Goldstein;Shinji Watanabe;G. Anumanchipalli
Jiachen Lian;A. Black;Yijingxiu Lu;L. Goldstein;Shinji Watanabe;G. Anumanchipalli
中科院分区:
其他
文献类型:
--
作者:
Jiachen Lian;A. Black;Yijingxiu Lu;L. Goldstein;Shinji Watanabe;G. Anumanchipalli

文献摘要

相似文献

发音表征学习是神经语音产生系统建模的基础研究。我们之前的工作已经建立了一个深入的范式,将发音运动学数据分解为手势,明确地模拟了人类语音产生机制编码的语音和语言结构,以及相应的手势分数。我们通过提出两个问题来继续这一系列工作:(1)在原始算法中,发音器纠缠在一起,使得一些发音器不利用有效的移动模式,这限制了手势和手势评分的可解释性;(2)EMA数据从发音器中稀疏采样,这限制了学习表示的可理解性。在这项工作中,我们提出了一种新的发音表示分解算法,该算法利用引导因子分析来获得特定于发音的因子和因子得分。然后对因子分数采用神经卷积矩阵因子分解算法以导出新的姿势和姿势分数。我们的实验与rtMRI语料库,捕捉细粒度的声道轮廓。主观和客观的评价结果表明,新提出的系统提供的发音表示,是可理解的,可推广的,有效的和可解释的。
Articulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly model the phonological and linguistic structure encoded with human speech production mechanism, and corresponding gestural scores. We continue with this line of work by raising two concerns: (1) The articulators are entangled together in the original algorithm such that some of the articulators do not leverage effective moving patterns, which limits the interpretability of both gestures and gestural scores; (2) The EMA data is sparsely sampled from articulators, which limits the intelligibility of learned representations. In this work, we propose a novel articulatory representation decomposition algorithm that takes the advantage of guided factor analysis to derive the articulatory-specific factors and factor scores. A neural convolutive matrix factorization algorithm is then employed on the factor scores to derive the new gestures and gestural scores. We experiment with the rtMRI corpus that captures the fine-grained vocal tract contours. Both subjective and objective evaluation results suggest that the newly proposed system delivers the articulatory representations that are intelligible, generalizable, efficient and interpretable.