Measuring Speech Activity

Measuring Speech Activity
复制标题

测量言语活动

DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
P. Kabal
P. Kabal
中科院分区:
--
文献类型:
--
作者:
P. Kabal

文献摘要

被引文献

相似文献

本报告讨论了ITU-T建议P.56中描述的用于测量活动语音电平的算法。P.56中的方法B确定语音活动因子,该语音活动因子表示信号被认为是活动语音(与背景空闲噪声相对)的时间分数以及信号的语音部分的相应活动电平。基本算法在每个采样时间生成包络值。将包络值与一组离散阈值进行比较。通过在阈值之间的对数域中插值来确定(近似)活动语音电平。在这份报告中,我们评估的语音活跃水平的影响,由于插值。建议P.56允许采样率低至600 Hz。子采样数据的结果进行了比较,在全语音采样率计算。测量语音活动1测量语音活动语音活动测量涉及确定信号包含活动语音的时间分数以及语音活动时的语音电平。语音活动的知识在语音信号测量中是重要的。对于语音数据库,重要的是要确保不适当的前导和拖尾非语音被删除,并根据峰值信号电平和活动语音电平适当缩放语音电平[1]。为了测试具有环境噪声的语音编码器,通过将记录的背景噪声添加到干净的语音段来创建人工测试信号。这种语音加噪声信号的信噪比被确定为语音的有效电平与记录噪声的均方根电平之比[1]。在语音编码界,相当多的研究工作正在花费在可变速率编码器或不连续传输系统上,这些编码器或不连续传输系统试图通过利用语音出现在谈话突发中的事实来节省平均比特率和/或功耗。这种技术的功效可以与言语活动测量进行比较。在ITU-T(国际电信联盟,电信标准化部门)建议P.56 [2]中给出了测量语音信号电平的规范,作为方法B。语音的活跃水平的测量考虑到语音可能包含嵌入式停顿的事实。实验表明,如果有350-400 ms或更大的间隙,听众会感觉到语音中的停顿[3]。如果这种间隙是由于短语之间的停顿或强调单词的停顿,它们被称为语法停顿。语法停顿和其他长间隙与闲置噪音不影响感知响度,不被视为积极的讲话。任何话语中固有的较小间隙被称为结构性停顿,并被视为活动语音片段的一部分。语音活动性算法的输出是语音活动性因子,其表示可以被认为是活动语音的信号的部分和信号的语音部分的对应活动语音电平。使用建议P.56中的算法实现语音电压表是ITU-T软件工具库[4][5]的一部分,在此称为ITU-T STL。所讨论的算法给出了整个话语的活动水平信息。测量语音活动2其他语音水平测量依赖于语音水平的即时指示,并且用于水平的实时指示(参见[2]中方法A的讨论)。一个例子是音量单位(VU)米经常看到专业和消费者音频设备。1包络计算语音活动算法计算语音信号的“包络”。这是语音样本值的幅度的双指数滤波,pi = gpi-1 +(1-g)|习|,qi = gqi−1 +(1− g)|Pi|. (1)从零初始条件开始计算包络qi。1参数g由求平均的时间常数确定,并设置为
This report discusses the algorithm described in ITU-T Recommendation P.56 for measuring the active speech level. Method B in P.56 determines a speech activity factor representing the fraction of time that the signal is considered to be active speech (as opposed to background idle noise) and the corresponding active level for the speech part of the signal. The basic algorithm generates an envelope value at each sample time. The envelope values are compared with a discrete set of thresholds. The (approximate) active speech level is determined by interpolating in the log domain between the threshold values. In this report we assess the effects on the speech active level due to interpolation. Recommendation P.56 allows for sampling rates as low as 600 Hz. Results for subsampled data are compared with those calculated at the full speech sampling rate. Measuring Speech Activity 1 Measuring Speech Activity Speech activity measurement involves determining the fraction of time that a signal contains active speech and the speech level while speech is active. Knowledge of the speech activity is important in speech signal measurements. For speech data bases, it is important to ensure that undue leading and trailing non-speech be excised and that the speech level be properly scaled based on the peak signal level and the active speech level [1]. For testing speech coders with environmental noise, artificial test signals are created by adding recorded background noise to clean speech segments. The signal-to-noise ratio for such speech-plus-noise signals is determined as the ratio of the active level for the speech to the rms level for the recorded noise [1]. In the speech coding community, considerable research effort is being expended on variable rate coders or discontinuous transmission systems that attempt to economize on average bit rate and/or power consumption by exploiting the fact the speech occurs in talk spurts. The efficacy of such techniques can be compared to speech activity measurements. Specifications for the measurement of the level of speech signals are given in ITU-T (International Telecommunication Union, Telecommunication Standardization Sector) Recommendation P.56 [2] as Method B. The measurement of the active level of speech takes into account the fact that speech may contain embedded pauses. Experiments have shown that listeners will perceive a pause in the speech if there is a gap of 350–400 ms or larger [3]. If such gaps are due to pauses between phrases or pauses to emphasize words, they are termed grammatical pauses. Grammatical pauses and other long gaps with idle noise do not affect the perceived loudness and are not counted as active speech. The smaller gaps inherent in any utterance are termed structural pauses and are counted as part of the active speech segment. The output of the speech activity algorithm is a speech activity factor representing the fraction of the signal that can be considered to be active speech and the corresponding active speech level for the speech part of the signal. An implementation of a Speech Voltmeter using the algorithm in Recommendation P.56 is part of the ITU-T Software Tools Library [4][5] referred to here as ITU-T STL. The algorithm under discussion presents a active level information for an utterance as a whole. Measuring Speech Activity 2 Other speech level measurements rely on an immediate indication of the speech level and are meant for a real-time indication of level (see the discussion of Method A in [2]). An example is the volume unit (VU) meter often seen on both professional and consumer audio equipment. 1 Envelope Calculation The speech activity algorithm calculates an “envelope” for the speech signal. This is a double exponential filtering of the magnitude of the speech sample values, pi = gpi−1 + (1− g)|xi|, qi = gqi−1 + (1− g)|pi|. (1) The envelope qi is calculated starting with zero initial conditions.1 The parameter g is determined by the time constant of the averaging and is set to