Speaker verification based on the fusion of speech acoustics and inverted articulatory signals.

Speaker verification based on the fusion of speech acoustics and inverted articulatory signals.
复制标题

基于语音声学和反发音信号融合的说话人验证

DOI:
10.1016/j.csl.2015.05.003
复制
发表时间:
2016-03
影响因子:
4.3
通讯作者:
Narayanan S
Narayanan S
中科院分区:
计算机科学3区
文献类型:
--
作者:
Li M;Kim J;Lammert A;Ghosh PK;Ramanarayanan V;Narayanan S

文献摘要

相似文献

我们提出了一种实用的、特征级别和分数级别的融合方法,将声学信息和估计发音信息结合起来,用于文本无关和文本相关的说话人确认。从实际应用的角度出发,研究了如何将动态发音信息与常规声学特征相结合来提高说话人确认性能。在与文本无关的说话人确认中,我们发现将从测量的语音产生数据中获得的发音特征与传统的Mel频率倒谱系数(MFCC)连接在一起可以显著提高性能。然而,由于直接测量发音数据在许多现实世界的应用中是不可行的,我们也试验了通过声学到发音倒置获得的估计发音特征。我们探索了特征级别和分数级别的融合方法,发现即使在估计发音特征的情况下,系统的整体性能也有显着提高。这样的性能提升可能是由于在估计的发音特征中嵌入了说话人之间的变化信息。由于发音动力学包含重要信息,我们在文本相关说话人验证中加入了倒置发音轨迹。我们证明,倒置发音特征引入的发音约束有助于拒绝错误密码试验,并在分数级融合后提高性能。我们分别在X射线微束数据库和RSR 2015数据库上评估了拟议的方法,用于上述两项任务。实验结果表明,两种说话人确认任务的错误率均降低了15%以上。
We propose a practical, feature-level and score-level fusion approach by combining acoustic and estimated articulatory information for both text independent and text dependent speaker verification. From a practical point of view, we study how to improve speaker verification performance by combining dynamic articulatory information with the conventional acoustic features. On text independent speaker verification, we find that concatenating articulatory features obtained from measured speech production data with conventional Mel-frequency cepstral coefficients (MFCCs) improves the performance dramatically. However, since directly measuring articulatory data is not feasible in many real world applications, we also experiment with estimated articulatory features obtained through acoustic-to-articulatory inversion. We explore both feature level and score level fusion methods and find that the overall system performance is significantly enhanced even with estimated articulatory features. Such a performance boost could be due to the inter-speaker variation information embedded in the estimated articulatory features. Since the dynamics of articulation contain important information, we included inverted articulatory trajectories in text dependent speaker verification. We demonstrate that the articulatory constraints introduced by inverted articulatory features help to reject wrong password trials and improve the performance after score level fusion. We evaluate the proposed methods on the X-ray Microbeam database and the RSR 2015 database, respectively, for the aforementioned two tasks. Experimental results show that we achieve more than 15% relative equal error rate reduction for both speaker verification tasks.