Neural network based speaker classification and verification systems with enhanced features

Neural network based speaker classification and verification systems with enhanced features
复制标题

具有增强功能的基于神经网络的说话人分类和验证系统

DOI:
10.1109/intellisys.2017.8324265
复制
发表时间:
2017
期刊:
2017 Intelligent Systems Conference (IntelliSys)
影响因子:
--
通讯作者:
A. Ganapathiraju
A. Ganapathiraju
中科院分区:
--
文献类型:
--
作者:
Zhenhao Ge;A. N. Iyer;S. Cheluvaraja;R. Sundaram;A. Ganapathiraju

文献摘要

被引文献

相似文献

本文提出了一种基于前馈神经网络的独立文本说话人分类和验证框架,这是两个相关的说话人识别系统。通过优化特征和模型训练,在分类中实现了100%的分类率和小于6%的等错误率(ERR),分别使用了大约1秒和5秒的数据。语音主动检测(VAD)比常规检测更严格的特征确保提取更强的浊音部分用于说话人识别,说话人水平均值和方差归一化有助于消除同一说话人样本之间的差异。这两种方法都被证明可以提高系统性能。在构建神经网络说话人分类器时,采用网格搜索优化网络结构参数,并采用动态约简正则化参数避免训练终止于局部极小值。它使培训以更低的成本走得更远。在说话人验证中,预测分数归一化(奖励峰值明显的说话人身份指标,惩罚得分高但竞争对手较多的弱指标)和针对说话人的阈值设置(显著降低ROC曲线上的ERR)提高了性能。这里使用采样率为8K的TIMIT语料库。对200名男性说话者进行分类性能训练和测试。其中的测试文件作为域内注册说话人,其余126名男性说话人的数据作为域外说话人,即冒名顶替者进行说话人验证。
This work presents a novel framework based on feed-forward neural network for text-independent speaker classification and verification, two related systems of speaker recognition. With optimized features and model training, it achieves 100% classification rate in classification and less than 6% Equal Error Rate (ERR), using merely about 1 second and 5 seconds of data respectively. Features with stricter Voice Active Detection (VAD) than the regular one for speech recognition ensure extracting stronger voiced portion for speaker recognition, speaker-level mean and variance normalization helps to eliminate the discrepancy between samples from the same speaker. Both are proven to improve the system performance. In building the neural network speaker classifier, the network structure parameters are optimized with grid search and dynamically reduced regularization parameters are used to avoid training terminated in local minimum. It enables the training goes further with lower cost. In speaker verification, performance is improved with prediction score normalization, which rewards the speaker identity indices with distinct peaks and penalizes the weak ones with high scores but more competitors, and speaker-specific thresholding, which significantly reduces ERR in the ROC curve. TIMIT corpus with 8K sampling rate is used here. First 200 male speakerwje used to train and test the classification performance. The testing files of them are used as in-domain registered speakers, while data from the remaining 126 male speakers are used as out-of-domain speakers, i.e. imposters in speaker verification.