Comparison of models using time-frequency features for speech classification

Comparison of models using time-frequency features for speech classification
复制标题

使用时频特征进行语音分类的模型比较

DOI:
10.1109/rivf.2006.1696427
复制
发表时间:
2006
期刊:
2006 International Conference onResearch, Innovation and Vision for the Future
影响因子:
--
通讯作者:
G. Kubin
G. Kubin
中科院分区:
--
文献类型:
--
作者:
T. V. Pham;G. Kubin

文献摘要

被引文献

相似文献

本文回顾和评价了两种不同的语音分类方法,使用多维特征来自二进小波变换(DyWT)。第一种方法是基于多阈值决策模型,而第二种方法是基于前馈神经网络的训练。将提取的特征与模型的自适应阈值进行比较,以根据声学类和语音组对输入语音信号帧进行分类。该方法进行了测试和评估TIMIT数据库在计算性别依赖。文中给出了语音分类器的构造过程,并将这两种算法与其它算法进行了性能比较,讨论了它们的优点和局限性。语音自动分类是许多语音处理方法和各种语音应用的关键。语音活动检测是语音分类的一个直接应用,被大多数语音应用所采用。在自动语音识别中,为了提高识别率,需要一种语音分类器来改善端点检测的性能。在拼接点选择合适的平滑策略可以提高拼接语音合成的性能。一些语音编码系统使用语音分类来确定每个不同语音帧的最佳比特分配。在互联网电话应用中,自适应丢失隐藏算法在发送端使用有声/无声检测器。这有助于接收器基于丢失段和相邻段之间的相似性来隐藏信息的丢失。此外,通过将语音分类器作为预分类步骤,可以更快地执行大型数据库的语音对齐。自20世纪80年代以来,许多文章通过多种方法研究了语音分类任务。原则上,通过依赖于从输入语音帧中提取的不同类型的特征向量来进行分类。这些特征可以通过三种方法导出:·第一种方法在时域中工作并使用统计测量。它们的共同特征是过零率、相对能级、自相关系数等(1)-(3)。这种方法只有在使用大量参数时才能达到良好的精度。
This paper reviews and evaluates two different ways of speech classification using multidimensional features derived from the Dyadic Wavelet Transform (DyWT). The first method is based on the Multi-Threshold Decision Model, while the second method is based on the training of Feedforward Neural Networks. The extracted features are compared against adaptive thresholds of the models to classify the input speech signal frames in terms of acoustic classes and phonetic groups. The methods are tested and assessed with the TIMIT database in counting gender dependency. The paper presents the procedures of building the speech classifiers, shows the performance comparison of the two algorithms with other algorithms, and discusses their advantages as well as limitations. Automatic speech classification is crucial for many speech processing methods and in various speech applications. Voice activity detection which is employed by most speech applications is a direct application of speech classification. In automatic speech recognition, there is a need for a phonetic classifier to improve the performance of endpoint detection in order to increase the recognition rate. The performance of concatenative speech synthesis may be improved by selecting proper smoothing strategies at concatenation points. Some speech coding systems use speech classification to determine the optimal bit allocation for every different speech frame. In internet telephony applications, the adaptive loss concealment algorithm uses the voiced/unvoiced detector at the sender. This helps the receiver to conceal the loss of information based on the similarity between the lost segments and the adjacent segments. Besides, the phonetic alignment of huge databases can be performed faster by applying the phonetic classifier as a pre-classification step. The speech classification task has been studied in many articles by a variety of methods since the 1980's. In principle, the classification is done by relying on different types of feature vectors which are extracted from the input speech frames. These features can be derived by three approaches: • The first approach works in the time domain and uses statistical measurements. The common features are zero crossing rate, relative energy level, autocorrelation coef- ficients, etc. (1)-(3). This approach only achieves good accuracy if using a large number of parameters.