Speaker Identification Using Pseudo Pitch Synchronized Phase Information in Voiced Sound

Speaker Identification Using Pseudo Pitch Synchronized Phase Information in Voiced Sound
复制标题

DOI:
--
复制
发表时间:
2011
期刊:
--
影响因子:
--
通讯作者:
Kohta Shimada;Kazumasa Yamamoto;S. Nakagawa
Kohta Shimada;Kazumasa Yamamoto;S. Nakagawa
中科院分区:
其他
文献类型:
--
作者:
Kohta Shimada;Kazumasa Yamamoto;S. Nakagawa

文献摘要

相似文献

在传统的基于Mel倒谱系数(MFCC)的说话人辨认方法中,忽略了相位信息。我们最近的研究表明,相位信息包含说话人相关的特征。我们提出了一种新的提取方法,只从浊音段提取基音同步相位信息。说话人识别实验使用NTT CLEAN数据库和JNAS数据库进行。使用新的相位提取方法,我们在两个数据库上分别获得了大约27%和46%的说话人错误率的相对降低。当我们将相位信息与基于MFCC的方法相结合时,我们也分别获得了大约52%和42%的相对误差减小。在传统的基于Mel倒谱系数(MFCC)的说话人识别方法中,只使用了语音帧中傅里叶变换的大小。这意味着忽略了相位分量。当然,MFCC不仅捕获特定于说话人的声道信息,还捕获声源特征。然而,从激励源特征提取的特征参数对于说话人识别也是有用的[1]、[4]、[5]、[6]、[7]、[10]。现有的几乎所有方法都是基于线性预测编码(LPC)分析的。马尔科夫和Nakagawa提出了一种基于高斯混合模型(GMM)的与文本无关的说话人识别系统,该系统将基音和LPC残差与LPC得到的倒谱系数结合在一起[4]。他们的实验结果表明,当考虑到基音和倒谱系数之间的相关性时,利用基音信息是最有效的。在文献[5]中,提出了一种自动估计和建模语音的声门流导数源波形并将模型参数应用于说话人识别的技术。与传统MFCC中的信息相比,残留阶段中说话人特定信息的互补性在[6]中得到了证明。通过线性预测分析,从语音信号中提取剩余相位。郑等人。提出了一种基于声源激励和声道系统的互补声学特征的说话人确认系统[7]。提出了一种新的特征集,称为残差小波倍频程系数(WOCOR),用于捕捉嵌入在线性预测残差信号中的谱时态源激励特征[7]。最近,许多基于群时延的相位信息的说话人识别研究已经被提出[8]、[9]。Wang等人。提出了用于说话人识别的相位相关特征[11]。这种类型的相位信息考虑了所有频率范围。我们认为,由于相位信息捕捉到了源信号的特征,因此对于说话人识别是有效的。以前,我们提出了一种结合MFCC和相位信息的说话人识别系统[1]、[2],直接从语音的傅立叶变换的有限带宽中提取。我们还证明了相位信息对于清洁和噪声环境下的说话人识别是有效的[1]、[2]、[3]。然而,由于加窗位置的影响,在提取相位信息时出现了一些问题。因此,我们提出了一种仅提取浊音基音同步相位信息的新方法。使用新的提取方法,在NTT和JNAS数据库上的说话人辨识率分别提高了约27%和46%。本文的其余部分组织如下。第二节介绍了相位信息提取方法,第三节讨论了相位和MFCC相结合的方法。第四节报告了实验装置和结果,第五节给出了我们的结论。公式[1],[3]通过对输入语音信号序列S(ω,t)=X(ω,t)+Jy(ω,t)=ωX2(√,t)+Y2(ω,t)×e进行离散傅立叶变换得到信号的频谱(ω,t)。然而,即使在相同频率ω下,相位也根据输入语音的限幅位置而变化。为了克服这一问题,保持某一基频ω的相位不变,并据此估计其他频率的相位。例如,通过将基频ω设置为π/4,我们得到S‘(ω,t)=√X2(ω,t)+Y2(ω,t)×e×Ej(π4−θ(ω,t),(2)而对于另一个频率ω’=2πf‘,频谱变为S’(ω‘,t)=√X2(ω’,t)+Y2(ω‘,t)×e’,t)×eω‘ω(π4−θ(ω,t)=X̃(ω’,T)+jỸ(ω‘,t)。(3)APSIPA ASC 2011西安
In conventional speaker identification methods based on mel-frequency cepstral coefficients (MFCCs), phase information is ignored. Our recent studies have shown that phase information contains speaker dependent characteristics. We propose a new extraction method to extract pitch synchronous phase information from the voiced section only. Speaker identification experiments were performed using the NTT clean database and JNAS database. Using the new phase extraction method, we obtained a relative reduction in the speaker error rate of approximately 27% and 46%, respectively, for the two databases. We also obtained a relative error reduction of approximately 52% and 42%, respectively, when combining phase information with the MFCC-based method. I. I NTRODUCTION In conventional speaker identification methods based on mel-frequency cepstral coefficients (MFCCs), only the magnitude of the Fourier Transform in time-domain speech frames is used. This means that the phase component is ignored. Of course, MFCCs capture not only speaker-specific vocal tract information, but also vocal source characteristics. Nevertheless, feature parameters extracted from excitation source characteristics are also useful for speaker identification [1], [4], [5], [6], [7], [10]. Almost all of the existing methods are based on Linear Predictive Coding (LPC) analysis. Markov and Nakagawa proposed a Gaussian Mixture Model (GMM) based text-independent speaker identification system that integrates pitch and the LPC residual with the LPC-derived cepstral coefficients [4]. Their experimental results show that using pitch information is the most effective when the correlation between pitch and the cepstral coefficients is taken into consideration. An automatic technique for estimating and modeling the glottal flow derivative source waveform of speech and applying the model parameters to speaker identification was proposed in [5]. The complementary nature of speakerspecific information in the residual phase compared with the information in conventional MFCCs was demonstrated in [6]. The residual phase was derived from speech signals by linear prediction analysis. Zheng et al. proposed a speaker verification system using complementary acoustic features derived from vocal source excitation and the vocal-tract system [7]. A new feature set, called the wavelet octave coefficients of residues (WOCOR), was proposed to capture the spectrotemporal source excitation characteristics embedded in the linear predictive residual signal [7]. Recently, many speaker recognition studies using group delay based phase information have been proposed [8], [9]. Wang et al. proposed phaserelated features for speaker recognition [11]. This type of phase information considers all frequency ranges. We think that phase information is valid for speaker identification, since it captures the features of the source wave. Previously, we proposed a speaker identification system using a combination of MFCCs and phase information [1], [2], directly extracted from the limited bandwidth of the Fourier transform of the speech wave. We also showed that the phase information is effective for speaker identification in clean and noisy environments [1], [2], [3]. However, problems occurred in extracting the phase information because of the influence of the windowing position. Therefore, we propose a new method to extract pitch synchronous phase information in voiced sound only. Using the new extraction method, the speaker identification rate improved by approximately 27% and 46% for the NTT and JNAS databases, respectively. The rest of this paper is organized as follows. Section 2 presents the phase information extraction method, while Section 3 discusses combining the phase and MFCC methods. The experimental setup and results are reported in Section 4, and Section 5 presents our conclusions. II. PHASE INFORMATION EXTRACTION A. Formulas [1], [3] The spectrumS(ω, t) of a signal is obtained by DFT of an input speech signal sequence S(ω, t) = X(ω, t) + jY (ω, t) = √ X2(ω, t) + Y 2(ω, t)× e. (1) However, the phase changes, depending on the clipping position of the input speech even at the same frequency ω. To overcome this problem, the phase of a certain basis frequency ω is kept constant, and the phases of other frequencies are estimated relative to this. For example, by setting the basis frequencyω to π/4, we obtain S′(ω, t) = √ X2(ω, t) + Y 2(ω, t)× e × ej(π4 −θ(ω,t)), (2) whereas for the other frequency ω′ = 2πf ′, the spectrum becomes S′(ω′, t) = √ X2(ω′, t) + Y 2(ω′, t)× e ′,t) × e ω ′ ω ( π 4 −θ(ω,t)) = X̃(ω′, t) + jỸ (ω′, t). (3) APSIPA ASC 2011 Xi’an